AI Security

Prompt Injection Explained: The Vulnerability Class Your Scanner Cannot See

Prompt injection is not a bug in a particular product. It is a structural consequence of systems that cannot distinguish instructions from data — and it is already in production across the region.

Researcher examining AI system behaviour

Every vulnerability class has a moment where it moves from academic curiosity to something attackers use routinely. Prompt injection reached that point while most organisations were still deciding whether to allow AI tools at all.

What it is

A language model receives a single stream of text. Some of it is your instruction; some of it is data the system fetched. The model has no reliable way to tell which is which. If an attacker controls any of that data, they can write text that reads as an instruction — and the model may follow it.

This is structurally the same problem as SQL injection, with one important difference. SQL injection is solved by parameterised queries: a clean separation between code and data. No equivalent separation exists for natural language, which is why this class has proven far harder to eliminate.

Where the untrusted text comes from

Review every path by which text reaches your model:

  • Web pages the system browses or summarises.
  • Emails, support tickets and chat messages it reads.
  • Documents, CVs and invoices it processes — including text hidden in white-on-white or at tiny font sizes.
  • Code and configuration it ingests from repositories.
  • Output from one tool that becomes input to the next, which is where multi-step agents get interesting.
  • Its own memory or retrieval store, if anything untrusted can be written there — an injection that persists is considerably worse than one that fires once.

What an attacker gets

In our assessments, successful injection typically achieves one of four things:

  • Data exfiltration — the agent is instructed to include sensitive context in a response, a tool call, or a URL it fetches.
  • Unauthorised tool use — it is persuaded to call something it should not, on the attacker's behalf.
  • Output manipulation — a CV instructs the screening agent to rank it highly; an invoice instructs the processing agent to approve it.
  • Guardrail bypass — the operational instructions are overridden, usually by convincing the model that the rules do not apply in this case.

Defences, in order of how much they help

Limit what the agent can do

Covered at length elsewhere, and still the highest-value control. If a manipulated agent cannot do anything consequential, injection becomes an annoyance rather than an incident.

Separate trust levels

Where the architecture allows it, do not let the component that reads untrusted content be the same component that holds credentials and calls tools. A summariser with no tool access that hands structured output to a privileged component is substantially harder to abuse than one agent doing both.

Constrain the output

If a step should produce one of five values, enforce that programmatically rather than trusting the model to comply. Structured output that is validated before use removes a large share of the practical attack surface.

Filter, but do not rely on it

Input and output filtering catches the obvious attempts and raises the effort required. It is a speed bump, not a boundary. Any defence that depends on recognising malicious natural language will eventually meet phrasing it does not recognise.

Test it properly

Scanners do not find this. It requires someone sitting with your actual deployment, with your actual tools connected, trying to make it misbehave — which is exactly what our AI agent assessments do.