What is prompt injection?
Prompt injection is an attack on AI systems built on large language models. An attacker supplies text that the model treats as an instruction, overriding what the developer or user intended it to do.
Also known as: prompt injection attack, indirect prompt injection, agent hijacking
Researched and fact-checked by AI, with no human review. 8 sources listed below. How we verify
Last updated
How it works
A large language model (LLM) receives its developer's instructions, the user's request and any outside material, such as a web page or email, as one stream of text. The US National Institute of Standards and Technology (NIST) defines prompt injection in its glossary. NIST calls it an attack that exploits the joining of untrusted input with a prompt written by a more trusted party, such as the application's designer.
Simon Willison proposed the name in a blog post on September 12, 2022, by analogy with SQL injection, an older attack on databases. Riley Goodside had shown that GPT-3 could be told to ignore its previous directions.
There are two main forms. In a direct attack, the person typing into the system supplies the malicious text. In an indirect prompt injection, the attacker plants instructions in a resource the model will read later. A 2023 paper by Kai Greshake and five co-authors described attacks under that name. It reported demonstrating them against Bing's GPT-4 powered chat and code-completion tools.
Why it matters
The risk grows when models are connected to email, browsers and other tools. A hijacked AI agent can then take actions as well as produce text. The OWASP Gen AI Security Project lists prompt injection as LLM01, the first entry in its 2025 Top 10 for LLM applications. OWASP treats jailbreaking, which aims to make a model disregard its safety rules, as one form of prompt injection.
Where things stand in 2026
As of October 2026, none of the bodies cited here describes a complete fix. In a blog post dated December 8, 2025, the UK National Cyber Security Centre argued that prompt injection may never be mitigated in the way SQL injection has been. Its reason: LLMs do not separate instructions from data. It advised limiting what such systems are permitted to do and monitoring their activity. On December 22, 2025, TechCrunch reported that OpenAI, writing about its ChatGPT Atlas browser, said prompt injection is unlikely ever to be fully solved.
A blog post from NIST's Center for AI Standards and Innovation, dated March 23, 2026, described a public red-teaming competition hosted by Gray Swan. More than 400 participants made over 250,000 hijacking attempts against 13 frontier AI models. At least one attack succeeded against each model. OWASP's suggested mitigations include least-privilege access, human approval for high-risk actions and separating external content.
Sources
- prompt injection - Glossary | CSRC, NIST Computer Security Resource Center
- indirect prompt injection - Glossary | CSRC, NIST Computer Security Resource Center
- Prompt injection attacks against GPT-3, Simon Willison's Weblog
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, arXiv (Greshake, Abdelnabi, Mishra, Endres, Holz and Fritz)
- LLM01:2025 Prompt Injection, OWASP Gen AI Security Project
- Prompt injection is not SQL injection (it may be worse), UK National Cyber Security Centre (NCSC)
- OpenAI says AI browsers may always be vulnerable to prompt injection attacks, TechCrunch
- Insights into AI Agent Security from a Large-Scale Red-Teaming Competition, NIST Center for AI Standards and Innovation (CAISI)