PROMPT INJECTION PROTECTION

Screen every prompt
before your model sees it.

Prompt injection is an attack where instructions hidden in content your AI reads, such as a web page, email or PDF, are followed by the model instead of yours. AiDren screens every prompt with a judge model before it reaches OpenAI, Anthropic or Mistral, blocks injection attempts, fails closed, and logs the reason and score.

No card required ~5ms proxy overhead One-line setup
HOW IT WORKS

How AiDren blocks prompt injection
A checkpoint on the only path in.

Your app already sends every LLM call to one endpoint. Point that endpoint at AiDren and each request is classified before it is forwarded — no SDK rewrite, no change to the request body.

  1. 1

    A request arrives

    Your app calls api.aidren.co.uk with an AiDren proxy key instead of your provider key. Nothing else about the request changes.

  2. 2

    The judge model screens it

    A fast classifier — with an automatic fallback provider — checks the request for injected or hidden instructions. Short, clearly-benign messages skip it; the rest add roughly 200–500 ms.

  3. 3

    Clean passes, dirty is blocked

    Clean traffic goes upstream untouched. An injection attempt is blocked with a reason and a score, logged on your Events page, and never reaches your model.

WHAT IT STOPS

Attacks it is built to catch

Real, shipped screening — not a keyword list you have to keep feeding.

  • Indirect prompt injection — instructions buried in content your agent was asked to read or summarise.
  • Direct jailbreaks — “ignore previous instructions” and its many rewrites.
  • Multi-turn attacks — payloads split across several messages; the built-in attack test exercises these.
  • Hidden instructions — text planted to be read by the model but not the user.
  • Fails closed — if screening cannot complete, the request is blocked, not forwarded.
  • Clean traffic untouched — no rewriting, no injected system prompts, no broken tool calls.
DEFENCE IN DEPTH

Where input screening fits

Screening is one layer. OWASP lists prompt injection as LLM01 in its 2025 Top 10 for LLM Applications and recommends several mitigations together: constrained model behaviour, input and output filtering, privilege control, human approval for high-risk actions, segregating untrusted content and adversarial testing. The NCSC says prompt injection “may never be totally mitigated in the way that SQL injection attacks can be”.

Defence layers against prompt injection
Defence layerWhat it helps withWhat it cannot doWhere AiDren fits
Strong system promptCasual misuseStop determined or indirect injectionTest yours with the free attack test
Input screening (judge model)Recognisable injection and jailbreak attempts, including multi-turnCatch every novel phrasingCore feature of AiDren
Output scanningLeaks of secrets, personal data and the system promptUndo actions already takenOutput scanning
Least privilege and human approvalLimits the damage a successful attack can doNot a filter, so it is your designEgress reporting flags unexpected outbound calls
IN THE DASHBOARD

Every decision, logged with its reason.

app.aidren.co.uk/events LIVE
AiDren Events dashboard showing clean, blocked, and flagged requests, each with the reason it was screened

A real screen from a live AiDren account.

FAQ

Questions,
answered plainly.

More in the API reference and on the pricing section.

What is prompt injection?

Prompt injection is when someone hides instructions inside content your AI reads — a webpage, an email, a file — hoping your model follows them instead of you. AiDren screens every request before it reaches your model, catching injected instructions before they can hijack your agent.

How much latency does AiDren add?

The proxy layer itself adds about 5 ms. On top of that, requests that need screening get one classification call — usually 200–500 ms, which runs before your request reaches OpenAI, Anthropic, or Mistral, so it overlaps nothing. Short, clearly-benign messages skip the classification call entirely. Model-file scans and egress reports add nothing to your chat traffic.

What if AiDren's screening is unavailable?

AiDren fails closed. If the classification pass can't complete (the screening provider is down, a timeout, an outage) the request is blocked, not waved through. We'd rather your app get an error it can retry than silently forward an unscreened request to your model. The screening layer also runs on a primary provider with an automatic fallback, so a single provider outage doesn't take it down.

Which LLM providers are supported?

OpenAI, Anthropic, and Mistral today, with the same drop-in proxy pattern for all three. More providers are coming.

Is there a way to see what AiDren would actually catch, before I commit to anything?

Yes, via the built-in attack test. Paste your system prompt, and AiDren runs it against roughly 20 real prompt-injection and jailbreak attempts on a cheap model using your own upstream key, then shows you exactly what got through with your AiDren protection on versus off. It's included on every plan, including the trial.

What is the difference between direct and indirect prompt injection?

Direct injection is a user typing instructions that override the model's rules, such as “ignore previous instructions”. Indirect injection hides the instructions in content the model is asked to process, like a web page, email or file. Indirect is usually harder to spot because the user never sees it. AiDren screens for both.

Can a classifier alone stop prompt injection?

No. A classifier reduces the attacks that reach your model, but the UK NCSC notes that there is no inherent distinction between data and instructions in an LLM. Pair screening with least-privilege tool access, output checks and human approval for high-impact actions.

Does it work for multi-turn attacks?

Yes. AiDren screens multi-turn attacks, where a payload is split across several messages. Requests that need screening add roughly 200 to 500 ms for the classification call, and short, clearly benign messages skip it.

RELATED GUIDES

Related guides

READY WHEN YOU ARE

Put a checkpoint between
your agent and its inputs.

Get the full stack free for 14 days. No card, no sales call, no SDK rewrite.

Start protecting requests