Guide
Glossary
Prompt injection
Prompt injection is an attack in which text an AI system reads contains instructions that override the ones it was given. The text can be typed straight into a chat, or planted in an email, web page or document the AI later processes, which is called indirect prompt injection. Language models cannot reliably tell trusted instructions from untrusted content, so a successful injection can make an assistant reveal information, give misleading answers or, in an AI agent, take actions nobody intended. OWASP lists prompt injection as the top risk for applications built on large language models. There is no complete fix, so systems are designed to limit the damage.
Also called indirect prompt injection, direct prompt injection
Last updated:
Example
In a law firm
A law firm's intake agent reads new enquiries. One emailed enquiry attaches a PDF with hidden text telling the AI to forward recent client correspondence to an outside address. Because the agent can only draft replies for staff approval and has no access to other matters, the instruction goes nowhere, and the attempt shows up in the audit log.
How can a firm reduce the risk of prompt injection?
A firm reduces prompt injection risk by limiting what its AI tools can reach and do, rather than relying on the model to spot attacks. Give each assistant or agent access only to the data its task needs, require human approval before anything is sent or changed, keep an audit log of every action, and treat any email, document or web page from outside the firm as untrusted input. The OWASP Top 10 for LLM applications sets out further controls.
Guides that explain it in context
Secure by design. Set up correctly. Fully managed.
Talk to us before you commit to anything
Start with a free 45-minute discovery call. We look at your systems and priorities, then recommend a first step with a fixed scope, or tell you if we are not the right fit.
