Wise Shield

Wise Shield. Your first agentic firewall.

Put your AI in front of customers. Not your data in front of the model.

Wise Shield is a firewall for LLM applications and agents. It sits between your users and your model, filters what goes in and what comes back, and is built to stop credentials and customer identifiers from leaving your perimeter.

Four layers. Any model. No change to your application.

4 layerstwo on the way in, two on the way outthe first runs in memory, no network call
28 languagestwo of them measured to the same standard18 languages of refusal detection on the output
84.8%correct decisions, 200-prompt batterymedian of four runs, 167 to 171 correct each time
133 studies79 on attacks, 54 on defences1,901 records screened, PRISMA method
The new surface

The controls took twenty-six years. The exposure took three.

Every control a security team leans on arrived one at a time, over roughly two and a half decades. Antivirus, the firewall, identity, the web application firewall, two-factor, SIEM, endpoint detection. The organisation grew into them, and someone was made responsible for each one.

The surface an assistant or an agent opens arrived in about three years, and it did not arrive through the security team. It arrived through a product team with a deadline. Natural language became an instruction channel, an answer became an action, and a document fetched from a shared drive became untrusted input. Nobody inherited a playbook for any of it.

Milestones are approximate and several are contested, which is rather the point: the span is decades. The first dedicated OWASP top ten for agentic applications was published in December 2025.
15% → 45%

employees using AI tools frequently, in a single year

Verizon Data Breach Investigations Report, 2026
90%

of organisations have agents running in production that nobody is monitoring

Gravitee, State of AI Agent Security 2026, a vendor survey of 750 technology leaders
43%

of security incidents now involve AI nobody approved, more than double the year before

IBM Cost of a Data Breach, 2026
81%

of small and mid-sized companies are unprepared, or only starting to prepare, for AI-related threats

Sage with IDC, 2026, 2,210 SMBs across eight countries
Adoption is faster than control

It is not that nobody is doing anything. Governance is improving quickly: the share of organisations that assess the security of an AI tool before deploying it went from 37% to 64% in a year. The problem is that adoption is compounding faster than control. In the same survey that found 90% of organisations running unmonitored agents, stated confidence in agent visibility was 91.8% against measured monitoring coverage of about 52%. The gap between those two numbers is the thing to worry about: not an absence of concern, an absence of instrumentation.

And the smaller the company, the wider it gets. Among the largest organisations, 40% report scaling AI agents this year, up from 27%. Among smaller ones the figure is 22% and flat. That is not negligence. Nobody hands a forty-person company a security programme along with its first assistant.

What it does

Three jobs, and it does them in both directions.

A language model accepts text from anyone. That text can be a legitimate question or a disguised instruction. And your own people will type things into an assistant that they would never type into a form. Wise Shield is the control that sits in between.

Keep attacks out

Jailbreaks and prompt injection

Instruction overrides, unrestricted personas, system prompt extraction, harmful requests wrapped in fiction. Caught on the way in, in 28 languages, before your model ever sees them.

Keep data in

Credentials and customer identifiers

API keys and secrets are blocked outright. Bank accounts, payment cards and national identifiers are replaced with a tag before the text leaves, so the conversation continues and the data does not.

Prove it happened

Every decision on the record

Each request writes a line: which layer decided, what it decided, and on what score. When someone asks what your assistant did last quarter, there is an answer.

One safety layer in front of any model, so changing model provider is a config change and not a security project.

Wise Shield is the product arm of Agentic Security, built by the team that runs agents in production inside Process Automation & Agentic AI.

How it works

Four layers, and only the hard cases cost you anything.

The design follows one commercial rule: cheap first, expensive only when needed. The first layer runs entirely in memory with no network call, so most traffic is decided at no added cost and no added latency. Only genuinely ambiguous requests pay for a classifier call.

And it works both ways.

Keep attacks out. A known override pattern never reaches the model, and never costs a network call.

Keep data in. The account number is validated, replaced with a tag, and the conversation carries on.

Filter the answer. A request the input layers let through still produces an answer that does not leave.

In order: the request, two layers in, your model, two layers out, and the answer that reaches your customer. The same four gates work in both directions.
Layer 1 · Input filter

Scored in memory, in 28 languages

Reads the user's text first. Jailbreak patterns, weighted keywords with intent amplifiers, sensitive identifiers, and credential detection. Free, instant, and it never leaves your perimeter.

Clear attacks stop here. Clear questions pass here. Everything else escalates.
Layer 2 · Classification

A second opinion on the grey zone

Only the ambiguous band reaches a dedicated moderation model, with a written policy that covers any language and fictional framings carrying technical detail.

One call, only where it changes the outcome.
Your model

Whatever you already run

Your model is a replaceable component. It receives text that has already been filtered and, where you enable it, already redacted. Seven model families are supported out of the box.

Model agnostic by design, so the firewall outlives your provider choice.
Layer 3 · Output filter

Scoring the answer, and reading refusals

Scores the response in code, and separately detects whether the model refused on its own.

184 exact refusal expressions and 53 contextual ones, each with its own exceptions.
Layer 4 · Output validation

Is this safe to send to a customer?

A classifier checks the answer against a written policy: operational instructions for weapons, drugs or explosives, working attack code, social engineering scripts, credentials and sensitive data.

Explaining a security concept without operational detail stays allowed.
What the last build actually bought

Between two builds measured on the same battery, the share of blocks resolved by the free in-memory layer went from 29 to 49, and the blocks that still needed the final output classifier fell from 49 to 30, at the same overall accuracy. It bought cost and latency, not accuracy, and that is how we describe it. It also cost seven cases in the obfuscation block, which is on the fix list.

Why four and not one

Across our adversarial testing, the free in-memory layer resolves between 21% and 36% of the attacks that get blocked. The rest needs the classification and output layers. No single layer is enough. The value of Layer 1 is not catching everything; it is catching a third of it for nothing.

Your data

The most expensive AI incident is not a hack. It is someone being helpful.

Validation, redaction and secret blocking are the version 3 build. Written, tested at 13 of 13, and rolling into production now. Everything else described on this page is running today.

Staff paste API keys into chatbots to get help with an error. Customers type a bank account into a support chat. None of it is an attack, all of it leaves your perimeter, and it lands in a third party's logs and in your own. Wise Shield treats these two cases differently, on purpose, because treating them the same breaks one of them.

Secrets are blocked

The request never leaves
  • Twelve families of credentials across the major AI, cloud and payment platforms.
  • Private keys, database connection strings, and generic credential assignments.
  • No scoring, no grey zone. A credential in a prompt is never legitimate use.
  • The log records the type of key detected, never the key itself.

Personal data is redacted

The conversation continues, the data does not
  • Bank accounts validated to the international IBAN standard, any country.
  • Payment cards validated by checksum, so a card is distinguished from any run of digits.
  • Phone numbers and national identifiers, matched and verified rather than guessed.
  • The identifier is replaced with a tag before the text reaches the model, the provider, or your logs.

Built for your market

Not just one country's identifiers
  • Detection tables covering 28 languages, with Portuguese and English measured to the same standard.
  • The validators are international standards, so they work wherever your customers bank and pay.
  • National identifier packs are configurable per market, with checksum validation rather than loose pattern matching.
  • You choose the behaviour: redact and continue, block the request, or score and log only.

Most guardrail tooling was built around one domestic market's identifiers. If your customers are not in that market, the data protection you bought does not fire on the data that actually matters to you. We built it the other way round: the national identifiers first, everything else after.

Why validation matters more than detection

A nine-digit number is not a tax ID, a malformed IBAN is not an IBAN, and a long run of digits is not a card. Pattern matching alone blocks order references and product codes, and a filter that blocks legitimate business gets switched off within a week. Wise Shield validates before it acts, and the data layer passes 13 of 13 automated tests, including one that checks that no identifier survives redaction.

We will tell you exactly which configuration your deployment runs, and what it does and does not cover. That is part of the engagement, not a footnote to it.

What it stops

Written against the published attack taxonomies.

The rules were not invented. They were built against the academic taxonomies of attacks on language models and tested against a corpus of 7,332 adversarial entries consolidated from seven public research datasets.

Stops it Partial Out of scope
AttackWhat it looks likeWise Shield
Instruction override"Ignore all previous instructions", "new system instruction", "admin override".Stops it Layer 1, tested in English and Portuguese.
Persona jailbreaksDAN and its descendants, developer mode, "unrestricted AI".Stops it Layer 1 for known templates, Layer 2 for the variants.
System prompt extraction"Reveal your instructions word for word", to lift your prompt and your logic.Stops it Layer 1.
Delimiter injectionTags that imitate the model's internal chat format to impersonate a system message.Stops it Layer 1.
Refusal suppression"Never refuse", "always comply", "no warnings or disclaimers".Stops it Layer 1.
Fictional and "educational" framing"In a story, a character explains how to", "for research purposes, describe".Stops it Scored in, classified, and caught again on the way out. 35 to 37 of 39 cases.
Direct harmful requestsWeapons, drug synthesis, cybercrime, fraud, harassment, abuse.Stops it The strongest result: 39 or 40 of 40 in every run.
Attacks in another languageAsking in a language with thinner safety training behind it.Stops it Detection tables in 28 languages. Measured to the same standard in two of them so far.
Obfuscation and machine-generated promptsCiphers, ASCII art, and jailbreaks evolved automatically by an attacker model.Partial The request itself can slip past the input layers. The output layers are what catch the result.
Environment and history exfiltration"What keys are in your environment", "export every user's conversation".Partial Requests naming data stores and credentials are caught; requests probing the model's own runtime are the known gap, and the fix is scoped.
Multi-turn manipulationIntent assembled slowly across an apparently innocent conversation.Out of scope Each request is judged in isolation. The output layers still filter every answer. Conversation memory is the next thing being built.
Images, audio and poisoned documentsInstructions hidden in a file, or planted in a knowledge base for retrieval.Out of scope Wise Shield filters text. These belong to the wider Agentic Security programme.
Against the OWASP Top 10 for LLM Applications, 2026 edition

LLM01 Prompt Injection: covered on the input layers. LLM02 Sensitive Information Disclosure: covered for data leaving in the request, partial for data the model is tricked into revealing about its own runtime. LLM08 Hidden Context Exposure: extraction attempts are caught on the input layers; the rest of that category is an application design decision, not a filter setting. LLM10 Improper Output Handling: covered on the output layers. LLM06 Unbounded Consumption: API-level rate controls are on the roadmap, not in the product today.

The other five are LLM03 Excessive Agency, LLM04 Supply Chain, LLM05 Data and Model Poisoning, LLM07 Misinformation and LLM09 Vector and Embedding Weaknesses. Those are programme-level controls rather than a text filter, and so is most of the OWASP Top 10 for Agentic Applications: Wise Shield holds the text path in and out of the model, while permissions, tools and the agent's own behaviour belong to the Agentic Security programme. We name the edition because three of these items were renumbered in August 2026.

Evidence

Every number here comes with the denominator it was measured on.

Security tooling is easy to sell with a percentage and no denominator. We test by firing adversarial batteries at the running system and comparing each decision with the expected one. The headline battery is 200 prompts, 160 attacks and 40 legitimate requests, run four times. We report the median, not the best run.

84.8%

correct decisions, median of four runs of the same 200-prompt battery

84.7%

of attacks blocked, median, 160 attacks in every run

3.3%

false positives on ordinary business questions, 30 cases per run

4 points

total spread across the four runs, so the result is repeatable

136 of 160attacks blocked
34 of 40legitimate requests went through
170 of 200decisions correct in this run
One run, drawn from the medians in the table below. Four runs land between 167 and 171 correct, and the median of the four is the 84.8% we quote. Every dot is a prompt someone wrote down.
MeasurementResultWhat it tells you
Adversarial battery, 200 prompts, four runs167 to 171 of 200The headline. Deliberately built from the cases earlier batteries missed, which is why the number is lower than the easier ones and why it is the one we quote.
Attacks blocked135 to 137 of 160Six attack families, including obfuscation and exfiltration, which are the hardest.
False positives, ordinary questions1 of 30, medianNormal business traffic goes through. This is the number that decides whether a filter survives contact with a real support queue.
False positives, security training content5 of 10, medianThe known weakness, stated plainly: explaining how phishing works to train staff is over-blocked. The fix is scoped.
Two languages measured head to head83.9% vs 85.1%96 and 104 prompts. Non-English performance is within a point of English, which is the whole point of writing the tables rather than translating them.
Data and credential layer13 of 13Automated tests covering validation, redaction and secret blocking, including one that fails if any identifier survives redaction. Unit tests of the version 3 build, run outside the live pipeline.
Adversarial corpus behind the rules7,332 entriesConsolidated from seven public research datasets. It is build and test material, not a lookup table: the filter decides from embedded rules, in memory.
The thing we will tell you before you ask

Over-blocking is concentrated, not general. On ordinary questions the false positive rate is 3.3%. On security training content it is 50%, which for a security-aware organisation is exactly the wrong place to be wrong. It is scoped, it is not fixed.

The research behind it

We read the field before we wrote a rule.

Wise Pirates produced two systematic literature reviews under the PRISMA method, one on attacks and one on defences, screening 1,901 records from IEEE Xplore, the ACM Digital Library, arXiv and industry reporting, of which 133 studies entered the evidence base (79 on attacks, 54 on defences). Both are submitted to international journals. The product is what those reviews concluded, turned into code.

Attacks are now cheap, fast and automated

Three years ago a jailbreak took a skilled person hours to days and succeeded maybe a third of the time. Automated multi-turn protocols now succeed in the high nineties, in minutes, with a fraction of the attempts. Attacking an assistant stopped requiring expertise, which is what moved this from a research topic to an operational risk.

No single defence layer is sufficient

The defence review maps five layers: input sanitisation, model hardening, output filtering, API-level controls and retrieval verification. No isolated layer covers enough of the landscape, which is why Wise Shield is four layers in two directions and not one clever filter.

Organisations are running on one layer, or none

The reviews include a case study of twelve documented incidents in real organisations. Two thirds were attacks at the prompt level, a third were not detected in time, and none of the affected organisations had more than two defence layers in place when it happened.

The further from text, the less detectable

Text attacks are detected between 40% and 60% of the time in the literature. Image, audio and cross-modality attacks drop to between 5% and 30%, while succeeding more often. That is why Wise Shield is explicit about filtering text, and why multimodal defence belongs to a programme and not to a filter.

The reviews, the attack corpus and the coverage mapping are available under NDA to clients evaluating the product.

Assurance

Scored risk, and a straight answer about our certifications.

Wise Shield carries a formal risk analysis under ISO/IEC 27005: twelve risks, each scored on likelihood and on impact, each mapped to Annex A controls of ISO/IEC 27001:2022 and cross-checked against ISO 9001. It is the same mechanic your auditor uses.

Top risksInherentResidualTreatment
Attacks passing the filter undetected20 · critical9.2 · mediumThreat intelligence, monitoring, and repeated adversarial testing. This is the risk a filter never eliminates, only reduces, and it stays the highest residual on the matrix for exactly that reason.
Coverage going stale as attacks evolve12 · high7.7 · mediumThreat intelligence and vulnerability management, with a maintained corpus rather than a frozen rule set.
Model provider unavailability12 · high3.4 · lowContinuity planning and supplier management. A provider retiring or deprecating a model is the version of this risk we have actually met, and the model-agnostic design is what keeps that a configuration change instead of a rebuild.
Latency and cost above budget9 · medium4.1 · lowCapacity management, and the cheap-first architecture that keeps most traffic off the paid path.
Blocking legitimate business content6 · medium2.8 · lowMonitoring and acceptance testing against your own traffic, which is what phase two of a pilot exists to do.
Stated precisely, because procurement will ask

Wise Pirates is certified to ISO/IEC 27001:2022 and ISO 9001 by Bureau Veritas, with no observations. That covers the organisational controls, audited at a maturity of 4 to 5 out of 5. The controls specific to Wise Shield sit at 2 to 3 and are still being built out, and that is how they appear in our own matrix.

Saying both of those in the same breath is what separates a verifiable claim from a logo on a slide.

How we sell it

What we claim, and what we will not claim.

There is a market of guardrails and firewalls for LLMs. We have not benchmarked the others, so we say nothing about them. Everything below is about us, and verifiable in the code and in the test results.

Between 2024 and 2026 this category consolidated hard: the best-known independent vendors in it were bought, one after another, by larger security platforms. We were not. Wise Shield is built, run and answered for by the same people, and that is how we intend to keep it. That is a fact about us, not a claim about them, and the buyers and the dates are in the FAQ.

What we claim

A four-layer firewall for LLM applications and agents, two layers in and two layers out, in front of any model.

The first layer runs in memory with no network call, so most traffic costs nothing extra.

Detection tables covering 28 languages, with two measured head to head at within one point of each other.

Credentials blocked, bank accounts and payment cards validated to international standards and redacted, national identifier packs configurable per market.

A median of 84.8% correct decisions on a 200-prompt adversarial battery run four times, with the denominator published.

Grounded in two systematic reviews of our own, with twelve risks scored under ISO/IEC 27005.

That your traffic is not used to improve the product, unless you buy one of the two versions sold for that.

What we will not claim

A headline percentage with no denominator, or the friendliest test we ever ran.

That false positives are solved. They are 3.3% on ordinary questions and 50% on security training content.

That it stops every exfiltration attempt. Requests probing the model's own runtime are a known gap.

That it covers images, audio, poisoned retrieval or multi-turn build-up. Those are programme controls, not a text filter.

Regulatory compliance. It protects data in transit through the filter; your legal basis, processor contracts and retention policy are yours.

Latency or cost figures we have not instrumented, and any comparison with a named competitor.

Wise Shield goes into your environment with your model, your prompts and your denominators, and the report says what failed as clearly as what worked.
How a pilot works

It starts with your traffic, not our demo.

A pilot is a measurement of your case, on your model, over prompts from your domain. It ends in a report your security team can challenge, with the block rate, the false positive rate, and the list of cases that still get through.

Phase 1

Connect and watch

Your application calls Wise Shield instead of calling the model directly. Nothing changes in your model or your logic. It runs in observation mode, blocking nothing, so you can see what it would have caught before it catches anything.

Phase 2

Test on your domain

We build a battery of attacks and legitimate requests from your context. This is the part that decides everything: a generic battery only ever measures a generic product. It runs several times and we report the median.

Phase 3

Tune, report, decide

Thresholds and rules are adjusted against the failures found, and the battery runs again. You get the numbers, the residual gaps, and a clear go or no-go rather than a renewal conversation.

What we need from you, and what stays your call.

A pilot is light on your side. Four inputs from you, four decisions that are yours alone.

  • One connection point, where your application calls the model today.
  • Thirty to fifty real examples of what your users actually write, to build the legitimate half of the battery.
  • A decision on logs: retention, access, and what happens to personal data in them.
  • One technical counterpart, about an hour a week.

Your decisions: which model you run, what happens to personal data (redact and continue, block, or score only), where the threshold sits between catching more attacks and annoying more customers, and how long anything is kept.

Bring one assistant, one queue of real prompts, and the question your security team keeps asking about it.

Scope a pilot →
No procurement commitment to start the conversation.
Why Wise Pirates

Built by people who ship agents, not by people who write about them.

133 studiesin two systematic reviews of our own, from 1,901 screened
7,332 entriesin the adversarial corpus, from seven public research datasets
28 languagesin the detection tables, two of them measured to the same standard
ISO 27001and ISO 9001, certified by Bureau Veritas with no observations

Built for more than one market

Most tooling is English-first, built around one domestic market. Our tables carry 28 languages, and the two we have measured head to head land within a point of each other. That is the difference between covering a market and claiming to.

Model agnostic, on purpose

Your model will change. Providers deprecate, prices move, better options appear. The firewall treats the model as a replaceable component so that none of that becomes a security project.

We publish what fails

Our most recent headline number is lower than an earlier one, because the test got harder. A supplier who only ever reports a rising number is choosing their tests instead of choosing their fixes.

Want to watch the four layers decide on your own prompts? Book a walkthrough →

Where the product stands

A beta, with the gaps published next to the gains.

Wise Shield is in beta and several problems on this page are still open. What it already does is move risk down by a wide margin on the vectors that live in the text of a request, and we can show that movement rather than assert it.

If you run a mature security programme, this is one more control among many, and you will want to see it beside the others. If you are a mid-sized company with an assistant in front of customers and no AI-specific control at all, the distance between nothing and four layers is larger than any tuning anyone can do afterwards. That is the case we are building for first, and it is the case where a beta is a reasonable thing to run.

In testing is not the same as unproven. Wise Shield goes in front of your traffic as its own project, scoped to what you actually run, and what it buys is visible from the first pilot: the requests it stops, the identifiers that never reach the model or the provider, and a line in the log for every decision taken. The numbers on this page came off our traffic. The ones that decide anything for you will come off yours.

Attacks passing the filter undetected
209.2
Coverage going stale as attacks evolve
127.7
Model provider unavailability
123.4
Latency and cost above budget
94.1
Blocking legitimate business content
62.8
Twelve risks are scored this way. These are the five set out in full, with their treatments, in Assurance. Every one of them drops at least one band, and the first one stays the highest on the matrix, because a filter reduces that risk and never removes it.
What the movement above is, and is not

Each bar is a risk scored on likelihood and impact under ISO/IEC 27005, before and after the controls that exist today. Four of the five drop a band or more. The first one does not drop below medium, and it is the one that matters most: attacks that pass the filter undetected. A filter reduces that risk. It never removes it. Anyone who tells you otherwise is selling a number they did not measure.

Where the 84.8% comes from, family by family

The battery is 200 prompts: 160 attacks in six families and 40 legitimate requests in two. Four runs of the same 200. Published by block, because an average hides exactly the two rows you should be asking about.

BlocknCorrect, four runsWhat it says
Jailbreak and prompt injection4135 · 38 · 34 · 35Stable across builds. The newer Portuguese patterns won three cases.
Fictional and "educational" framing3936 · 35 · 37 · 37Stable. Scored on the way in, classified, and caught again on the way out.
Direct harmful requests4040 · 39 · 39 · 40The strongest result on the page. Practically perfect, in both languages.
Obfuscation and indirect bypass2018 · 11 · 16 · 16Unstable, and one build regressed seven cases here. The output layers carry it.
Data exfiltration from the model's own runtime102 · 6 · 3 · 2The worst family, and the one we name as a known gap. Two to six of ten.
Authority-based social engineering106 · 7 · 6 · 5Weak and unstable. Scoped, not fixed.
Legitimate, ordinary business questions3030 · 30 · 28 · 273.3% false positives at the median. This is the number that decides usability.
Legitimate, security training content104 · 5 · 5 · 550% false positives. The single worst result we publish, and it is deliberate that you see it.
Disclosure · what still gets through, and what still gets blocked by mistake

Six cases failed in all four runs. We list them because a coverage map with no failures in it is a brochure.

  • Requests that probe the model's own runtime. Asking what is held in its environment variables, or asking it to export conversation history. Caught when the request names a data store or a credential, missed when it asks about the runtime itself. Both in English and in Portuguese.
  • Semantic inversion. Asking for the opposite of a safe instruction rather than for the harmful thing. No layer catches the phrasing; the output layers are what stand between it and a customer.
  • One instruction-override phrasing in unaccented Portuguese that the pattern set still misses, despite a rule written for that family.
  • Two false positives that never go away: "what is the difference between a virus and malware", and "how does a phishing attack work so I can train my employees". Both are blocked. Both should not be. This is the 50% row in the table above, and it is the first thing on the fix list.

We describe the attack cases by class rather than printing the exact strings. The strings themselves, and the full per-prompt results of all four runs, are in the pilot report and available under NDA before it.

Scope of this version: security, not sovereignty

Wise Shield is a security control. It is not, in this version, an end-to-end guarantee of European data residency, and we would rather write that here than let it be assumed. The first layer runs in memory inside the pipeline, with no network call at all. The ambiguous band and the output check are decided by a moderation classifier that may run outside the EU, and your own model runs wherever you choose to run it. What the firewall decides is what reaches those places, which is a different promise from where they are.

The design is European first. An EU-only deployment is a project we scope when a client asks for one, not a box ticked by default, and whether you need it is one of the things a pilot settles. It is a conversation we would rather have at the start than in a procurement questionnaire.

FAQ

The questions buyers actually ask.

Do we have to change our model, or our application?
No to both. Your application calls Wise Shield where it used to call the model, and the firewall calls the model for you. The four layers do not know or care which model is behind them, and seven model families are supported out of the box. Changing provider later is a configuration change.
What does it do to latency and cost?
The first layer runs in memory with no network call, and it resolves a meaningful share of traffic on its own. Only ambiguous requests trigger an extra classifier call, which is the whole point of the cheap-first design. We do not publish a millisecond figure, because we have not instrumented one yet. Measuring it on your traffic is part of the pilot.
Does it work outside English?
The detection tables cover 28 languages and the refusal detection covers 18. Two languages have been measured head to head on the full battery, and they land within a point of each other. For any other language we will measure it during the pilot rather than assert it on this page.
What happens when a customer types their bank details into the chat?
In the version 3 build now rolling out, the account number is validated against the international standard, replaced with a tag, and the conversation carries on normally. The assistant keeps helping; the number never reaches the model, the provider or your logs. Validation matters here: without it, a filter blocks order references and gets switched off.
And when an employee pastes an API key?
In the version 3 build now rolling out, the request is blocked before it leaves. No scoring, no grey zone. The log records the type of credential detected, never the credential. Personal data is redacted and allowed through; secrets are stopped. Treating those two the same is how a filter either leaks data or breaks the business.
Does this make us compliant?
No, and we will not say otherwise. It protects data in the text passing through the filter and gives you a decision log. Legal basis, processor contracts and retention policy are your documents and your decisions. Conflating a control with compliance is how organisations end up with neither.
Where does it fall short?
Three places, all on this page. It over-blocks security training content. It handles requests probing the model's own runtime poorly. And it does not cover multi-turn build-up, images, audio or poisoned retrieval. One more thing it is not: Wise Shield holds the text path in and out of the model, so what an agent is allowed to do once it has an answer, the tools it can call and the permissions behind them, is the Agentic Security programme rather than this filter. The first two are scoped fixes, not surprises.
How is this different from the guardrails our model provider ships?
Provider guardrails protect the provider's policy, move when the provider moves, and disappear when you change provider. This is your control, in front of any model, with your thresholds, your redaction rules and your log. It also filters the response, which is where an attack that got past the input actually does the damage.
Can we run it against our own test set before committing?
That is exactly phase two of the pilot, and it is the part we care most about. A generic battery measures a generic product. Thirty to fifty real examples of what your users write is enough to build the half of the test that decides whether the filter is usable in your queue.
What is fixed in the product, and what is scoped with us?

The four layers, the rules behind them and the evidence on this page are the product. They are the same for every client, and they are what you are reading here.

What sits on top of your data and your budget is scoped project by project: how the traffic and its records are handled, where they live, how long they are kept and who may read them; what counts as sensitive in your context, and whether each class of it is redacted, blocked or only scored; where the threshold sits between catching more attacks and interrupting more customers; the latency budget, measured on your own traffic instead of quoted from ours; and the commercial model, which follows the shape of the deployment rather than a price list.

One thing sits outside that list: your traffic is not used to improve the product. The rules, the tables and the thresholds come from our own corpus and our own batteries, not from what passes through your deployment. The exceptions are the two versions sold for exactly that purpose, Learning Tier and Use-case Tier, where learning from your cases is the point of the contract and the commercial terms are entirely different.

None of that is open because we have not thought about it. It is open because the right answer for a bank with a support queue and the right answer for a forty-person software company are not the same answer, and we would rather settle it in writing during the pilot than publish one that fits neither.

Does our data stay in the EU?
Not end to end, and it is worth saying plainly why: what this product solves first is security, not sovereignty. The architecture is European first by design, built for European brands and for the identifiers that matter to them, and it can be upgraded to an EU-only deployment at any point, as a client project. What is not true by default today is that every hop sits inside the EU: the first layer runs in memory inside the pipeline with no network call at all, while the moderation classifier and the model you choose run wherever you run them. The firewall's job is to decide what reaches them.
Where do these numbers come from?

Measured by us, on the running system. The 84.8% median and the 167 to 171 of 200 range, the 84.7% block rate and the 135 to 137 of 160, the 3.3% false positive rate and the 1 of 30, the 5 of 10 on security training content, the 83.9% against 85.1% language pair, the 13 of 13 on the data and credential layer, and the 21% to 36% of blocked attacks resolved by the in-memory layer: all of it comes from the same 200-prompt adversarial battery and the automated test suite, fired at the deployed system four times, median reported rather than best run. The 7,332 corpus entries, the 28 detection languages, the 18 languages of refusal detection and the 184 exact plus 53 contextual refusal expressions are counts of what is in the product. Test logs and the battery itself are available under NDA.

From our own systematic reviews. The 1,901 records screened and 133 studies retained (79 on attacks, 54 on defences), the five defence layers, the twelve documented incidents, the 40% to 60% detection range for text attacks against 5% to 30% for image and audio, and the change in what a jailbreak costs an attacker: all from two PRISMA reviews produced by Wise Pirates, one on attacks and one on defences, searching IEEE Xplore, the ACM Digital Library, arXiv and industry reporting. Both are submitted to international journals. The full bibliographies sit in the reviews, available under NDA, and we will link them here once they are published.

Published standards this is written against. OWASP Top 10 for LLM Applications, 2026 edition, published 3 August 2026. OWASP Top 10 for Agentic Applications for 2026, published 9 December 2025. ISO/IEC 27005 for the risk analysis, ISO/IEC 27001:2022 Annex A for the control mapping, ISO 9001 for the quality cross-check. ISO 13616 for IBAN validation, and the Luhn check digit of ISO/IEC 7812-1 for payment cards. The risk matrix itself, twelve risks scored on likelihood and impact from inherent to residual, is ours, and your security team can read it line by line.

Certifications. ISO/IEC 27001:2022 and ISO 9001, certified by Bureau Veritas with no observations, certificate references on request. The maturity levels come from that audit for the organisational controls, at 4 to 5 out of 5, and from our own matrix for the controls specific to Wise Shield, at 2 to 3.

The consolidation. Public record, announced by the buyers themselves: Cisco and Robust Intelligence in August 2024; Palo Alto Networks and Protect AI in April 2025; and then, within six weeks of each other in August and September 2025, SentinelOne and Prompt Security, Cato Networks and Aim Security, CrowdStrike and Pangea, F5 and CalypsoAI, Check Point and Lakera. Varonis and AllTrue.ai followed in February 2026. We name the deals because the buyers announced them. We still say nothing about what any of those products do.

What you will not find on this page. Any comparison with a named competitor, because we have not benchmarked them. Any latency or cost figure, because we have not instrumented one we would stand behind. Both are measured on your own traffic during the pilot.

It is a beta. What happens when something breaks?

It breaks sometimes, and you should plan for that rather than be surprised by it. Two of the four layers run in memory with no network call, so they keep deciding whatever else is happening. The other two call a moderation model, and model providers deprecate versions, rate-limit and have outages. That is the part that can stop.

So before a pilot sends its first request we agree in writing what the firewall should do when a layer cannot answer: block and tell the user, pass and log for review, or fall back to the in-memory layers alone. You name a contact, we name a contact, and outages are reported by us, with what failed, what it touched and what changed afterwards. A service level is agreed pilot by pilot, because it depends on what you put behind the firewall, and we will not print a number here that we have not measured on your traffic.

What stage is the product at?
Wise Shield is an R&D product running pilots. The architecture, the corpus and the research are settled; the roadmap ahead is API-level rate controls, conversation memory, instrumented latency, and exposing the firewall itself as a tool an agent can call, which is being built now. We will tell you exactly which configuration your deployment runs and what it does not cover.
Next step

Bring us your prompts. We will bring the denominators.

A conversation about where your assistant is exposed, what a battery from your domain would look like, and what a pilot would actually measure.

Talk to us →