Screen every prompt
before your model sees it.
Prompt injection is an attack where instructions hidden in content your AI reads, such as a web page, email or PDF, are followed by the model instead of yours. AiDren screens every prompt with a judge model before it reaches OpenAI, Anthropic or Mistral, blocks injection attempts, fails closed, and logs the reason and score.
How AiDren blocks prompt injection
A checkpoint on the only path in.
Your app already sends every LLM call to one endpoint. Point that endpoint at AiDren and each request is classified before it is forwarded — no SDK rewrite, no change to the request body.
-
1
A request arrives
Your app calls
api.aidren.co.ukwith an AiDren proxy key instead of your provider key. Nothing else about the request changes. -
2
The judge model screens it
A fast classifier — with an automatic fallback provider — checks the request for injected or hidden instructions. Short, clearly-benign messages skip it; the rest add roughly 200–500 ms.
-
3
Clean passes, dirty is blocked
Clean traffic goes upstream untouched. An injection attempt is blocked with a reason and a score, logged on your Events page, and never reaches your model.
Attacks it is built to catch
Real, shipped screening — not a keyword list you have to keep feeding.
- Indirect prompt injection — instructions buried in content your agent was asked to read or summarise.
- Direct jailbreaks — “ignore previous instructions” and its many rewrites.
- Multi-turn attacks — payloads split across several messages; the built-in attack test exercises these.
- Hidden instructions — text planted to be read by the model but not the user.
- Fails closed — if screening cannot complete, the request is blocked, not forwarded.
- Clean traffic untouched — no rewriting, no injected system prompts, no broken tool calls.
Related: Output & data-leak scanning · Pre-deploy attack test · AiDren vs Lakera · How AiDren compares →
Where input screening fits
Screening is one layer. OWASP lists prompt injection as LLM01 in its 2025 Top 10 for LLM Applications and recommends several mitigations together: constrained model behaviour, input and output filtering, privilege control, human approval for high-risk actions, segregating untrusted content and adversarial testing. The NCSC says prompt injection “may never be totally mitigated in the way that SQL injection attacks can be”.
| Defence layer | What it helps with | What it cannot do | Where AiDren fits |
|---|---|---|---|
| Strong system prompt | Casual misuse | Stop determined or indirect injection | Test yours with the free attack test |
| Input screening (judge model) | Recognisable injection and jailbreak attempts, including multi-turn | Catch every novel phrasing | Core feature of AiDren |
| Output scanning | Leaks of secrets, personal data and the system prompt | Undo actions already taken | Output scanning |
| Least privilege and human approval | Limits the damage a successful attack can do | Not a filter, so it is your design | Egress reporting flags unexpected outbound calls |
Every decision, logged with its reason.
A real screen from a live AiDren account.
What is prompt injection?
Prompt injection is when someone hides instructions inside content your AI reads — a webpage, an email, a file — hoping your model follows them instead of you. AiDren screens every request before it reaches your model, catching injected instructions before they can hijack your agent.
How much latency does AiDren add?
The proxy layer itself adds about 5 ms. On top of that, requests that need screening get one classification call — usually 200–500 ms, which runs before your request reaches OpenAI, Anthropic, or Mistral, so it overlaps nothing. Short, clearly-benign messages skip the classification call entirely. Model-file scans and egress reports add nothing to your chat traffic.
What if AiDren's screening is unavailable?
AiDren fails closed. If the classification pass can't complete (the screening provider is down, a timeout, an outage) the request is blocked, not waved through. We'd rather your app get an error it can retry than silently forward an unscreened request to your model. The screening layer also runs on a primary provider with an automatic fallback, so a single provider outage doesn't take it down.
Which LLM providers are supported?
OpenAI, Anthropic, and Mistral today, with the same drop-in proxy pattern for all three. More providers are coming.
Is there a way to see what AiDren would actually catch, before I commit to anything?
Yes, via the built-in attack test. Paste your system prompt, and AiDren runs it against roughly 20 real prompt-injection and jailbreak attempts on a cheap model using your own upstream key, then shows you exactly what got through with your AiDren protection on versus off. It's included on every plan, including the trial.
What is the difference between direct and indirect prompt injection?
Direct injection is a user typing instructions that override the model's rules, such as “ignore previous instructions”. Indirect injection hides the instructions in content the model is asked to process, like a web page, email or file. Indirect is usually harder to spot because the user never sees it. AiDren screens for both.
Can a classifier alone stop prompt injection?
No. A classifier reduces the attacks that reach your model, but the UK NCSC notes that there is no inherent distinction between data and instructions in an LLM. Pair screening with least-privilege tool access, output checks and human approval for high-impact actions.
Does it work for multi-turn attacks?
Yes. AiDren screens multi-turn attacks, where a payload is split across several messages. Requests that need screening add roughly 200 to 500 ms for the classification call, and short, clearly benign messages skip it.
Related guides
- What is prompt injection? Direct vs indirect, jailbreaks and the OWASP LLM01 mapping.
- How to prevent prompt injection A defence-in-depth checklist with code.
- Prompt injection examples Twelve attack patterns with sanitised payloads.
Put a checkpoint between
your agent and its inputs.
Get the full stack free for 14 days. No card, no sales call, no SDK rewrite.
Start protecting requests