Glossary // Prompt injection
What Is Prompt Injection? Definition, Examples and Defenses (2026)
Updated
Prompt injection is an attack where text the model reads as data, such as a web page, an email, a PDF, or a code comment, contains instructions that the model follows as if they came from you. It works because a language model has no hard boundary between instructions and content: your system prompt, the user's message, and a fetched web page all arrive as tokens in one context window, and a sufficiently convincing sentence in any of them can redirect the task.
The term was coined by Simon Willison in September 2022, days after Riley Goodside showed that a GPT-3 translation app could be told to "ignore the above directions" and say something else entirely. Willison named it after SQL injection on purpose: in both cases, untrusted input is concatenated into something that gets executed. The difference is that SQL has a fix (parameterised queries) and language models do not yet have an equivalent.
Direct and indirect prompt injection
- Direct injection is the user typing the attack into the chat box. It matters when the application has secrets or capabilities the user should not reach, for example a support bot whose system prompt contains internal policy, or a tool the user is not meant to trigger.
- Indirect injection is the attacker planting instructions in content the model will read later, without ever talking to it. A 2023 paper by Greshake and colleagues showed it against LLM-integrated apps, and it is the form that matters now, because agents read untrusted content all day: web pages, repository files, issue comments, inbound email, tool results.
The damage scales with what the model can do after reading. A chatbot that gets injected says something wrong. An agent with a shell, a browser, or an API token can delete files, send data out, or make changes that look like normal work.
Real examples from coding agents
Our news desk has covered a steady run of these in 2026, and the pattern repeats:
- A researcher showed that asking Claude Code to summarize a web page was enough to redirect it, using nothing more than hidden text on the page.
- GitSpawn demonstrated that coding agents will execute code from a repository they were only asked to clone and inspect.
- GitHub's AI agent leaked private repository contents when asked conversationally, an authorization failure that injection makes exploitable at scale.
- Work on secrets extraction from black-box LLMs showed that anything placed in a system prompt or context window should be treated as readable by a determined user.
None of these needed a novel exploit. They needed an agent that read untrusted content while holding real permissions.
Why filters and detectors miss it
The intuitive fix is to scan inputs for injection phrases. It does not hold up. An evaluation of open-source prompt injection detectors found they miss a meaningful share of realistic agent attacks, because a malicious instruction can be phrased as ordinary prose, split across documents, encoded, or written in another language. A detector trained on yesterday's payloads is a speed bump.
A more useful framing comes from research that treats prompt injection as a role problem, not a text problem: the model fails because it cannot verify who is entitled to give which instruction. Until models have that notion built in, the fix has to live outside the model.
The defenses that hold
Simon Willison's "lethal trifecta" is the most practical test. An agent is exposed when it combines three things: access to private data, exposure to untrusted content, and a way to send data out (a web request, an email, a commit, even a rendered image URL). Remove any one leg and most injection attacks lose their payoff. In practice:
- Least privilege per task. Scope tokens, directories, and tools to the job. A summarizer does not need write access; a code reviewer does not need network egress.
- Separate reading from acting. Have a restricted model read untrusted content and return structured fields, and let only those fields reach the agent that holds permissions. This is the idea behind Willison's dual-LLM pattern and Google DeepMind's CaMeL design.
- Mark untrusted content as data. Wrap fetched text in clearly labelled tags and tell the model that nothing inside them is an instruction. It is not a guarantee, but it raises the bar, and model providers now train for instruction hierarchy (system above user above tool output).
- Require confirmation for irreversible actions. Deleting, sending, paying, publishing, and pushing should need a human or a separate policy check, especially right after the agent read external content.
- Log and diff what the agent did after it read something. An unexpected tool call right after a fetch is the clearest injection signal you will get. Defenders are even planting honeypot injections in bait documents so a rogue agent reveals itself.
A prompt for auditing your own pipeline
You are a security reviewer for an LLM application.
<architecture>[describe each model call: what it reads, what tools and credentials it holds, what it can send out]</architecture>
For every model call, answer: 1) Which inputs could an outsider influence? 2) What is the worst action the call could take if those inputs contained instructions? 3) Does the call combine private data, untrusted content, and an outbound channel? Then rank the calls by risk and propose the smallest change that removes one leg of that combination for each of the top three.
It will not find every hole, but it forces the question most teams skip: what can this model do right after it reads something we did not write?
Prompt injection also has a long tail into memory. If an agent writes what it read into a persistent store, an injection can outlive the session that delivered it, which is why agent memory needs the same trust rules as the context window, whichever of the fragmenting memory architectures you pick.
Prompt injection in the news
- Open-Source Prompt Injection Detectors Struggle Against Realistic Agent Attacks
Sep 24, 2026
- Claude Code's Prompt Injection Problem Is Simpler Than You Think
Sep 1, 2026
- The Agent Infrastructure Stack Is Crystallizing: Sandboxes, Benchmarks, and Injection Attacks
Jul 19, 2026
- Defensive Prompt Injection: How Honeypot Prompts Catch Rogue AI Agents
Jul 14, 2026
- GitHub's AI Agent Leaks Private Repos via Prompt Injection
Jul 9, 2026
- Claude Code Is Becoming an Ecosystem: Memory, Workflow, and Injection Defense
Jul 6, 2026
- Prompt Injection Is a Role Problem, Not a Text Problem
Jun 23, 2026
- GitSpawn Shows AI Coding Agents Will Execute Whatever Code They Clone
Sep 5, 2026
READY TO ASCEND
Get prompt-engineering insight by email
New terms, prompt packs, and the news that changes how you prompt. No noise.
Questions
What is prompt injection in simple terms?
It is getting a language model to follow instructions it should have treated as data. The model reads your instructions and the content it is working on as one stream of text, so a sentence like 'ignore your previous instructions and do this instead' inside a web page or email can take over the task.
What is the difference between prompt injection and a jailbreak?
A jailbreak is a user talking a model out of its own safety rules. Prompt injection is a third party hijacking an application by planting instructions in content the application feeds the model. Jailbreaks embarrass the model provider; prompt injection attacks the developer's users and data.
What is indirect prompt injection?
It is injection where the attacker never talks to the model. They plant the instructions somewhere the model will read later: a web page an agent browses, a document in a RAG index, an email an assistant summarizes, a code comment a coding agent opens. It is the form that matters for agents.
Can prompt injection be fully prevented?
Not with current models. Filters and detectors reduce it, and instruction hierarchies make models more resistant, but none is reliable against an attacker who adapts. The dependable defenses are architectural: limit what a model that has read untrusted content is allowed to do.
Is prompt injection on the OWASP Top 10?
Yes. It is LLM01, the first entry, in the OWASP Top 10 for Large Language Model Applications, and has held that position since the list was first published.
Related terms
Glossary
Agent memory
Agent memory is the information an AI agent stores outside its context window, in files, databases, or indexes, and reads back later so it can remember across steps and sessions.
Glossary
Context compaction
Context compaction is when an AI agent nears its context window limit and replaces older conversation turns with a model-written summary so the session can keep going.
Prompt packs for this
Coding // Claude
Best Claude Prompts for Coding (2026)
13 Claude prompts for coding in 2026: debugging, refactoring, adversarial review, failing-test-first fixes, flaky tests, and instructions for Claude Code agent runs.
Agent Reliability // Claude
Silent Agent Failures: 8 Failure Modes and 10 Prompts to Catch Them (2026)
Silent agent failures (silent failures in AI agents) are runs that end in success while the work is wrong. The 8 failure modes behind them, the trace checks that catch them, and 10 prompts for reviewing agent transcripts.
Customer Support // ChatGPT
ChatGPT Prompts for Customer Support: 14 Reply, Macro and Escalation Templates (2026)
14 ChatGPT prompts for customer support: reply drafts, macros, de-escalation, ticket triage, refund explanations, escalation summaries, help docs from solved tickets, and a weekly ticket-theme report.