Glossary · Security & Resilience

What is prompt injection?

Short answer

Prompt injection is an attack on AI applications where text written by an attacker, typed in by a user or hidden in a web page, email or document the AI reads, contains instructions that override the developer’s intended behaviour. It can make an AI leak data, ignore its rules or misuse the tools it has access to.

Direct and indirect injection

  • Direct: the user types “ignore your previous instructions and…”. Annoying, but limited to what that user could already do.
  • Indirect: the instructions arrive inside content the AI processes on someone else’s behalf, such as a support email, a web page an agent browses, or a document in a RAG index. This is the dangerous case, especially for agents that can send email, call APIs or read private data.

Why it is hard to fix

An LLM reads instructions and data as the same stream of text, so there is no reliable way to make it ignore instructions inside data. Defences reduce the impact rather than prevent the attack.

How to limit the damage

  • Give the AI the least privilege it needs; treat it like an untrusted user.
  • Require human confirmation for actions with side effects.
  • Never let one step both read untrusted content and send data to an arbitrary destination.
  • Validate tool arguments like any other input, and allow-list URLs and recipients.
  • Log tool calls and review unusual ones.

Frequently asked questions

Can prompt injection be fully prevented?

Not today. Filters and careful prompts catch some attacks, but the reliable defence is limiting what a compromised AI can do: narrow permissions, human approval for risky actions and validated tool inputs.

Published · Updated · By · All terms

Go deeper