An LLM firewall
between your app and the model.
An LLM firewall is a security layer between your application and a language model that inspects prompts and responses and blocks prompt injection, data leaks and policy violations. Unlike a network firewall it judges meaning, not ports, so its decisions are probabilistic. AiDren is one: a drop-in proxy you adopt by changing a base URL.
How an LLM firewall works
Every call your application makes to a model passes through the firewall, in both directions. It decides whether to let the request or response through, change it, or block it.
- Request in. Your app sends a prompt, with its system prompt, conversation history and any retrieved content.
- Screen the prompt. A classifier, rules or both look for prompt injection, jailbreaks and policy violations. Clean requests are forwarded unchanged. Suspicious ones are blocked or flagged.
- Model responds. The upstream provider generates a reply.
- Scan the response. The firewall checks for leaked secrets, personal data and echoes of your system prompt, then passes, redacts or blocks.
- Log the decision. The outcome and reason are recorded so you can review them.
Because the checks are about meaning, they are statistical. That makes an LLM firewall useful and fallible in the same breath: it will catch a large share of recognisable attacks, and a determined attacker may still find a phrasing it misses. The NCSC's December 2025 guidance says prompt injection may never be totally mitigated, so treat the firewall as one layer. The prevention checklist shows the others.
LLM firewall vs WAF vs API gateway vs guardrails library
| WAF | API gateway | Guardrails library | LLM firewall | |
|---|---|---|---|---|
| Inspects | HTTP requests for web attack signatures | API calls: auth, routing, quotas | Prompts and outputs, from inside your code | Prompts and responses in transit |
| Understands language? | No | No | Yes | Yes |
| Where it runs | Edge or network | Network | In your application | Between app and model |
| Prompt injection | Not designed for it | Not designed for it | Depends on rules you configure | Core purpose |
| Integration effort | Low | Low to medium | Code changes | Often a base URL change |
These are complementary. Keep your WAF and gateway for what they do well. A guardrails library gives you fine control inside your code and means you maintain the integration. A firewall in front of the provider protects every call regardless of which service made it. Several products blur these categories, so judge each one by what it actually inspects.
What to look for in an LLM firewall
- Fail-closed behaviour. If screening cannot complete, does the request get blocked or waved through?
- Latency you can measure. Ask what a screened request adds, and whether clearly benign messages skip the expensive check.
- Both directions. Prompt screening without response scanning leaves the leak path open.
- Streaming support. Most chat products stream, and scanning must work on streamed output.
- Multi-turn awareness. Attacks are often split across messages.
- Per-key policy. Different apps need different rules, and a way to test a rule before enforcing it.
- Monitor mode. You should be able to see what it would block before it blocks anything.
- Logs with reasons. Every decision should be explainable and exportable.
- Honest limits. Be wary of any vendor that promises complete protection.
- Pricing you can read. Public prices let you compare without a sales call.
How AiDren implements an LLM firewall
AiDren is a proxy. You point your OpenAI-, Anthropic- or Mistral-compatible client at https://api.aidren.co.uk and use an AiDren proxy key. The request body is not rewritten. The API reference has the exact base URLs for each provider.
- Judge-model screening. A fast classifier, with an automatic fallback provider, checks requests for injected or hidden instructions, including multi-turn attacks. Short, clearly benign messages skip it. The proxy adds about 5 ms, and requests that need classification add roughly 200 to 500 ms. If classification cannot complete, the request is blocked. See prompt injection protection.
- Output scanning. Responses are checked for system-prompt leaks, emails, card numbers, IBANs, phone numbers, API keys and private keys, then redacted or blocked per key. Streaming responses are handled too, with the trade-offs described on the output scanning page. See output scanning.
- Custom policy. Term, regex, topic and threshold rules attach to a key and can be shadow-tested on live traffic first. See custom policy per key.
- Egress monitoring. An npm package reports unexpected outbound connections from an agent. It reports and never blocks. See egress monitoring.
- Model-file scanning. Downloads from Hugging Face Hub and GitHub are scanned for malicious code before they reach disk. See model-file scanning.
Monitor mode lets you watch what it would have blocked before enforcing anything, and you can try it against your own prompt with the free attack test.
What an LLM firewall cannot do
- It cannot guarantee that every attack is caught. Novel phrasings, encodings and indirect injection can get through any classifier.
- It cannot fix your permissions. If your agent can send money, a bypassed firewall means it can send money. Least privilege and human approval are design work.
- It does not cover risks outside the request path, such as poisoned training data, weak retrieval access control or hallucination. The OWASP LLM Top 10 guide shows which risks a proxy can touch.
- AiDren supports OpenAI, Anthropic and Mistral today. If you use another provider, it will not sit in that path.
Self-hosted libraries, enterprise platforms and AiDren
Broadly, the options fall into three groups. Open-source libraries that you host and tune yourself give maximum control and put the integration, hosting and upkeep on you. Enterprise platforms sold through sales teams bundle wider coverage and support, usually with a procurement process. Self-serve proxies like AiDren trade some of that breadth for a one-line integration, public prices and a trial you can start without talking to anyone.
Which fits depends on your team, your data-handling requirements and your budget. For sourced, dated detail on specific products, read the comparison pages: AiDren vs LLM Guard, AiDren vs Lakera and AiDren vs Protect AI, or the overview.
Public pricing
AiDren costs £49, £149 or £399 a month. Every plan includes all the layers above and unlimited team seats. The plan sets your monthly fair-use allowance of checked requests, not which features you can use. There is a 14-day free trial with no card, and billing is month to month. See the pricing page for allowances and annual terms.
What is an LLM firewall?
An LLM firewall is a layer between your application and a language model that inspects prompts and responses in real time to block prompt injection, data leakage and policy violations. Unlike a network firewall it analyses meaning rather than packets, so its decisions are probabilistic.
Is an LLM firewall the same as a WAF?
No. A web application firewall inspects HTTP traffic for known web attack signatures such as SQL injection and cross-site scripting. An LLM firewall examines natural-language prompts and model output, which have no fixed grammar. They solve different problems and can run together.
Can an LLM firewall stop prompt injection completely?
No. The UK NCSC says prompt injection may never be totally mitigated in the way SQL injection can be. An LLM firewall reduces how many attacks reach your model and catches some leaks on the way out, but you should also limit what the model can do and require approval for high-impact actions.
How do I add an LLM firewall to my app?
With a proxy-style firewall such as AiDren you change the API base URL to the proxy and use a proxy key in place of your provider key. The request and response format stays the same. Library-style tools are instead called from your code before and after each model call.
How much does AiDren cost?
Plans are £49, £149 and £399 a month, all with every feature, differing only in monthly request allowance, with a 14-day free trial and no card. Full details are on the pricing page.
Put a firewall in front
of your model.
Get the full stack free for 14 days. No card, no sales call, no SDK rewrite.
Start protecting requests