Wise Shield. Your first agentic firewall.
Put your AI in front of customers. Not your data in front of the model.
Wise Shield is a firewall for LLM applications and agents. It sits between your users and your model, filters what goes in and what comes back, and is built to stop credentials and customer identifiers from leaving your perimeter.
Four layers. Any model. No change to your application.
The controls took twenty-six years. The exposure took three.
Every control a security team leans on arrived one at a time, over roughly two and a half decades. Antivirus, the firewall, identity, the web application firewall, two-factor, SIEM, endpoint detection. The organisation grew into them, and someone was made responsible for each one.
The surface an assistant or an agent opens arrived in about three years, and it did not arrive through the security team. It arrived through a product team with a deadline. Natural language became an instruction channel, an answer became an action, and a document fetched from a shared drive became untrusted input. Nobody inherited a playbook for any of it.
employees using AI tools frequently, in a single year
Verizon Data Breach Investigations Report, 2026of organisations have agents running in production that nobody is monitoring
Gravitee, State of AI Agent Security 2026, a vendor survey of 750 technology leadersof security incidents now involve AI nobody approved, more than double the year before
IBM Cost of a Data Breach, 2026of small and mid-sized companies are unprepared, or only starting to prepare, for AI-related threats
Sage with IDC, 2026, 2,210 SMBs across eight countriesIt is not that nobody is doing anything. Governance is improving quickly: the share of organisations that assess the security of an AI tool before deploying it went from 37% to 64% in a year. The problem is that adoption is compounding faster than control. In the same survey that found 90% of organisations running unmonitored agents, stated confidence in agent visibility was 91.8% against measured monitoring coverage of about 52%. The gap between those two numbers is the thing to worry about: not an absence of concern, an absence of instrumentation.
And the smaller the company, the wider it gets. Among the largest organisations, 40% report scaling AI agents this year, up from 27%. Among smaller ones the figure is 22% and flat. That is not negligence. Nobody hands a forty-person company a security programme along with its first assistant.
Three jobs, and it does them in both directions.
A language model accepts text from anyone. That text can be a legitimate question or a disguised instruction. And your own people will type things into an assistant that they would never type into a form. Wise Shield is the control that sits in between.
Jailbreaks and prompt injection
Instruction overrides, unrestricted personas, system prompt extraction, harmful requests wrapped in fiction. Caught on the way in, in 28 languages, before your model ever sees them.
Credentials and customer identifiers
API keys and secrets are blocked outright. Bank accounts, payment cards and national identifiers are replaced with a tag before the text leaves, so the conversation continues and the data does not.
Every decision on the record
Each request writes a line: which layer decided, what it decided, and on what score. When someone asks what your assistant did last quarter, there is an answer.
Wise Shield is the product arm of Agentic Security, built by the team that runs agents in production inside Process Automation & Agentic AI.
Four layers, and only the hard cases cost you anything.
The design follows one commercial rule: cheap first, expensive only when needed. The first layer runs entirely in memory with no network call, so most traffic is decided at no added cost and no added latency. Only genuinely ambiguous requests pay for a classifier call.
And it works both ways.
Keep attacks out. A known override pattern never reaches the model, and never costs a network call.
Keep data in. The account number is validated, replaced with a tag, and the conversation carries on.
Filter the answer. A request the input layers let through still produces an answer that does not leave.
Scored in memory, in 28 languages
Reads the user's text first. Jailbreak patterns, weighted keywords with intent amplifiers, sensitive identifiers, and credential detection. Free, instant, and it never leaves your perimeter.
A second opinion on the grey zone
Only the ambiguous band reaches a dedicated moderation model, with a written policy that covers any language and fictional framings carrying technical detail.
Whatever you already run
Your model is a replaceable component. It receives text that has already been filtered and, where you enable it, already redacted. Seven model families are supported out of the box.
Scoring the answer, and reading refusals
Scores the response in code, and separately detects whether the model refused on its own.
Is this safe to send to a customer?
A classifier checks the answer against a written policy: operational instructions for weapons, drugs or explosives, working attack code, social engineering scripts, credentials and sensitive data.
Between two builds measured on the same battery, the share of blocks resolved by the free in-memory layer went from 29 to 49, and the blocks that still needed the final output classifier fell from 49 to 30, at the same overall accuracy. It bought cost and latency, not accuracy, and that is how we describe it. It also cost seven cases in the obfuscation block, which is on the fix list.
Across our adversarial testing, the free in-memory layer resolves between 21% and 36% of the attacks that get blocked. The rest needs the classification and output layers. No single layer is enough. The value of Layer 1 is not catching everything; it is catching a third of it for nothing.
The most expensive AI incident is not a hack. It is someone being helpful.
Validation, redaction and secret blocking are the version 3 build. Written, tested at 13 of 13, and rolling into production now. Everything else described on this page is running today.
Staff paste API keys into chatbots to get help with an error. Customers type a bank account into a support chat. None of it is an attack, all of it leaves your perimeter, and it lands in a third party's logs and in your own. Wise Shield treats these two cases differently, on purpose, because treating them the same breaks one of them.
Secrets are blocked
- Twelve families of credentials across the major AI, cloud and payment platforms.
- Private keys, database connection strings, and generic credential assignments.
- No scoring, no grey zone. A credential in a prompt is never legitimate use.
- The log records the type of key detected, never the key itself.
Personal data is redacted
- Bank accounts validated to the international IBAN standard, any country.
- Payment cards validated by checksum, so a card is distinguished from any run of digits.
- Phone numbers and national identifiers, matched and verified rather than guessed.
- The identifier is replaced with a tag before the text reaches the model, the provider, or your logs.
Built for your market
- Detection tables covering 28 languages, with Portuguese and English measured to the same standard.
- The validators are international standards, so they work wherever your customers bank and pay.
- National identifier packs are configurable per market, with checksum validation rather than loose pattern matching.
- You choose the behaviour: redact and continue, block the request, or score and log only.
Most guardrail tooling was built around one domestic market's identifiers. If your customers are not in that market, the data protection you bought does not fire on the data that actually matters to you. We built it the other way round: the national identifiers first, everything else after.
A nine-digit number is not a tax ID, a malformed IBAN is not an IBAN, and a long run of digits is not a card. Pattern matching alone blocks order references and product codes, and a filter that blocks legitimate business gets switched off within a week. Wise Shield validates before it acts, and the data layer passes 13 of 13 automated tests, including one that checks that no identifier survives redaction.
We will tell you exactly which configuration your deployment runs, and what it does and does not cover. That is part of the engagement, not a footnote to it.
Written against the published attack taxonomies.
The rules were not invented. They were built against the academic taxonomies of attacks on language models and tested against a corpus of 7,332 adversarial entries consolidated from seven public research datasets.
| Attack | What it looks like | Wise Shield |
|---|---|---|
| Instruction override | "Ignore all previous instructions", "new system instruction", "admin override". | Stops it Layer 1, tested in English and Portuguese. |
| Persona jailbreaks | DAN and its descendants, developer mode, "unrestricted AI". | Stops it Layer 1 for known templates, Layer 2 for the variants. |
| System prompt extraction | "Reveal your instructions word for word", to lift your prompt and your logic. | Stops it Layer 1. |
| Delimiter injection | Tags that imitate the model's internal chat format to impersonate a system message. | Stops it Layer 1. |
| Refusal suppression | "Never refuse", "always comply", "no warnings or disclaimers". | Stops it Layer 1. |
| Fictional and "educational" framing | "In a story, a character explains how to", "for research purposes, describe". | Stops it Scored in, classified, and caught again on the way out. 35 to 37 of 39 cases. |
| Direct harmful requests | Weapons, drug synthesis, cybercrime, fraud, harassment, abuse. | Stops it The strongest result: 39 or 40 of 40 in every run. |
| Attacks in another language | Asking in a language with thinner safety training behind it. | Stops it Detection tables in 28 languages. Measured to the same standard in two of them so far. |
| Obfuscation and machine-generated prompts | Ciphers, ASCII art, and jailbreaks evolved automatically by an attacker model. | Partial The request itself can slip past the input layers. The output layers are what catch the result. |
| Environment and history exfiltration | "What keys are in your environment", "export every user's conversation". | Partial Requests naming data stores and credentials are caught; requests probing the model's own runtime are the known gap, and the fix is scoped. |
| Multi-turn manipulation | Intent assembled slowly across an apparently innocent conversation. | Out of scope Each request is judged in isolation. The output layers still filter every answer. Conversation memory is the next thing being built. |
| Images, audio and poisoned documents | Instructions hidden in a file, or planted in a knowledge base for retrieval. | Out of scope Wise Shield filters text. These belong to the wider Agentic Security programme. |
LLM01 Prompt Injection: covered on the input layers. LLM02 Sensitive Information Disclosure: covered for data leaving in the request, partial for data the model is tricked into revealing about its own runtime. LLM08 Hidden Context Exposure: extraction attempts are caught on the input layers; the rest of that category is an application design decision, not a filter setting. LLM10 Improper Output Handling: covered on the output layers. LLM06 Unbounded Consumption: API-level rate controls are on the roadmap, not in the product today.
The other five are LLM03 Excessive Agency, LLM04 Supply Chain, LLM05 Data and Model Poisoning, LLM07 Misinformation and LLM09 Vector and Embedding Weaknesses. Those are programme-level controls rather than a text filter, and so is most of the OWASP Top 10 for Agentic Applications: Wise Shield holds the text path in and out of the model, while permissions, tools and the agent's own behaviour belong to the Agentic Security programme. We name the edition because three of these items were renumbered in August 2026.
Every number here comes with the denominator it was measured on.
Security tooling is easy to sell with a percentage and no denominator. We test by firing adversarial batteries at the running system and comparing each decision with the expected one. The headline battery is 200 prompts, 160 attacks and 40 legitimate requests, run four times. We report the median, not the best run.
correct decisions, median of four runs of the same 200-prompt battery
of attacks blocked, median, 160 attacks in every run
false positives on ordinary business questions, 30 cases per run
total spread across the four runs, so the result is repeatable
| Measurement | Result | What it tells you |
|---|---|---|
| Adversarial battery, 200 prompts, four runs | 167 to 171 of 200 | The headline. Deliberately built from the cases earlier batteries missed, which is why the number is lower than the easier ones and why it is the one we quote. |
| Attacks blocked | 135 to 137 of 160 | Six attack families, including obfuscation and exfiltration, which are the hardest. |
| False positives, ordinary questions | 1 of 30, median | Normal business traffic goes through. This is the number that decides whether a filter survives contact with a real support queue. |
| False positives, security training content | 5 of 10, median | The known weakness, stated plainly: explaining how phishing works to train staff is over-blocked. The fix is scoped. |
| Two languages measured head to head | 83.9% vs 85.1% | 96 and 104 prompts. Non-English performance is within a point of English, which is the whole point of writing the tables rather than translating them. |
| Data and credential layer | 13 of 13 | Automated tests covering validation, redaction and secret blocking, including one that fails if any identifier survives redaction. Unit tests of the version 3 build, run outside the live pipeline. |
| Adversarial corpus behind the rules | 7,332 entries | Consolidated from seven public research datasets. It is build and test material, not a lookup table: the filter decides from embedded rules, in memory. |
Over-blocking is concentrated, not general. On ordinary questions the false positive rate is 3.3%. On security training content it is 50%, which for a security-aware organisation is exactly the wrong place to be wrong. It is scoped, it is not fixed.
We read the field before we wrote a rule.
Wise Pirates produced two systematic literature reviews under the PRISMA method, one on attacks and one on defences, screening 1,901 records from IEEE Xplore, the ACM Digital Library, arXiv and industry reporting, of which 133 studies entered the evidence base (79 on attacks, 54 on defences). Both are submitted to international journals. The product is what those reviews concluded, turned into code.
Attacks are now cheap, fast and automated
Three years ago a jailbreak took a skilled person hours to days and succeeded maybe a third of the time. Automated multi-turn protocols now succeed in the high nineties, in minutes, with a fraction of the attempts. Attacking an assistant stopped requiring expertise, which is what moved this from a research topic to an operational risk.
No single defence layer is sufficient
The defence review maps five layers: input sanitisation, model hardening, output filtering, API-level controls and retrieval verification. No isolated layer covers enough of the landscape, which is why Wise Shield is four layers in two directions and not one clever filter.
Organisations are running on one layer, or none
The reviews include a case study of twelve documented incidents in real organisations. Two thirds were attacks at the prompt level, a third were not detected in time, and none of the affected organisations had more than two defence layers in place when it happened.
The further from text, the less detectable
Text attacks are detected between 40% and 60% of the time in the literature. Image, audio and cross-modality attacks drop to between 5% and 30%, while succeeding more often. That is why Wise Shield is explicit about filtering text, and why multimodal defence belongs to a programme and not to a filter.
The reviews, the attack corpus and the coverage mapping are available under NDA to clients evaluating the product.
Scored risk, and a straight answer about our certifications.
Wise Shield carries a formal risk analysis under ISO/IEC 27005: twelve risks, each scored on likelihood and on impact, each mapped to Annex A controls of ISO/IEC 27001:2022 and cross-checked against ISO 9001. It is the same mechanic your auditor uses.
| Top risks | Inherent | Residual | Treatment |
|---|---|---|---|
| Attacks passing the filter undetected | 20 · critical | 9.2 · medium | Threat intelligence, monitoring, and repeated adversarial testing. This is the risk a filter never eliminates, only reduces, and it stays the highest residual on the matrix for exactly that reason. |
| Coverage going stale as attacks evolve | 12 · high | 7.7 · medium | Threat intelligence and vulnerability management, with a maintained corpus rather than a frozen rule set. |
| Model provider unavailability | 12 · high | 3.4 · low | Continuity planning and supplier management. A provider retiring or deprecating a model is the version of this risk we have actually met, and the model-agnostic design is what keeps that a configuration change instead of a rebuild. |
| Latency and cost above budget | 9 · medium | 4.1 · low | Capacity management, and the cheap-first architecture that keeps most traffic off the paid path. |
| Blocking legitimate business content | 6 · medium | 2.8 · low | Monitoring and acceptance testing against your own traffic, which is what phase two of a pilot exists to do. |
Wise Pirates is certified to ISO/IEC 27001:2022 and ISO 9001 by Bureau Veritas, with no observations. That covers the organisational controls, audited at a maturity of 4 to 5 out of 5. The controls specific to Wise Shield sit at 2 to 3 and are still being built out, and that is how they appear in our own matrix.
Saying both of those in the same breath is what separates a verifiable claim from a logo on a slide.
What we claim, and what we will not claim.
There is a market of guardrails and firewalls for LLMs. We have not benchmarked the others, so we say nothing about them. Everything below is about us, and verifiable in the code and in the test results.
Between 2024 and 2026 this category consolidated hard: the best-known independent vendors in it were bought, one after another, by larger security platforms. We were not. Wise Shield is built, run and answered for by the same people, and that is how we intend to keep it. That is a fact about us, not a claim about them, and the buyers and the dates are in the FAQ.
What we claim
A four-layer firewall for LLM applications and agents, two layers in and two layers out, in front of any model.
The first layer runs in memory with no network call, so most traffic costs nothing extra.
Detection tables covering 28 languages, with two measured head to head at within one point of each other.
Credentials blocked, bank accounts and payment cards validated to international standards and redacted, national identifier packs configurable per market.
A median of 84.8% correct decisions on a 200-prompt adversarial battery run four times, with the denominator published.
Grounded in two systematic reviews of our own, with twelve risks scored under ISO/IEC 27005.
That your traffic is not used to improve the product, unless you buy one of the two versions sold for that.
What we will not claim
A headline percentage with no denominator, or the friendliest test we ever ran.
That false positives are solved. They are 3.3% on ordinary questions and 50% on security training content.
That it stops every exfiltration attempt. Requests probing the model's own runtime are a known gap.
That it covers images, audio, poisoned retrieval or multi-turn build-up. Those are programme controls, not a text filter.
Regulatory compliance. It protects data in transit through the filter; your legal basis, processor contracts and retention policy are yours.
Latency or cost figures we have not instrumented, and any comparison with a named competitor.
It starts with your traffic, not our demo.
A pilot is a measurement of your case, on your model, over prompts from your domain. It ends in a report your security team can challenge, with the block rate, the false positive rate, and the list of cases that still get through.
Connect and watch
Your application calls Wise Shield instead of calling the model directly. Nothing changes in your model or your logic. It runs in observation mode, blocking nothing, so you can see what it would have caught before it catches anything.
Test on your domain
We build a battery of attacks and legitimate requests from your context. This is the part that decides everything: a generic battery only ever measures a generic product. It runs several times and we report the median.
Tune, report, decide
Thresholds and rules are adjusted against the failures found, and the battery runs again. You get the numbers, the residual gaps, and a clear go or no-go rather than a renewal conversation.
What we need from you, and what stays your call.
A pilot is light on your side. Four inputs from you, four decisions that are yours alone.
- One connection point, where your application calls the model today.
- Thirty to fifty real examples of what your users actually write, to build the legitimate half of the battery.
- A decision on logs: retention, access, and what happens to personal data in them.
- One technical counterpart, about an hour a week.
Your decisions: which model you run, what happens to personal data (redact and continue, block, or score only), where the threshold sits between catching more attacks and annoying more customers, and how long anything is kept.
Bring one assistant, one queue of real prompts, and the question your security team keeps asking about it.
Scope a pilot →A filter is one layer. The programme around it is the rest.
Wise Shield covers input sanitisation and output filtering. Identity, tool supply chain, red teaming and continuous monitoring live next door.
Built by people who ship agents, not by people who write about them.
Built for more than one market
Most tooling is English-first, built around one domestic market. Our tables carry 28 languages, and the two we have measured head to head land within a point of each other. That is the difference between covering a market and claiming to.
Model agnostic, on purpose
Your model will change. Providers deprecate, prices move, better options appear. The firewall treats the model as a replaceable component so that none of that becomes a security project.
We publish what fails
Our most recent headline number is lower than an earlier one, because the test got harder. A supplier who only ever reports a rising number is choosing their tests instead of choosing their fixes.
Want to watch the four layers decide on your own prompts? Book a walkthrough →
A beta, with the gaps published next to the gains.
Wise Shield is in beta and several problems on this page are still open. What it already does is move risk down by a wide margin on the vectors that live in the text of a request, and we can show that movement rather than assert it.
If you run a mature security programme, this is one more control among many, and you will want to see it beside the others. If you are a mid-sized company with an assistant in front of customers and no AI-specific control at all, the distance between nothing and four layers is larger than any tuning anyone can do afterwards. That is the case we are building for first, and it is the case where a beta is a reasonable thing to run.
In testing is not the same as unproven. Wise Shield goes in front of your traffic as its own project, scoped to what you actually run, and what it buys is visible from the first pilot: the requests it stops, the identifiers that never reach the model or the provider, and a line in the log for every decision taken. The numbers on this page came off our traffic. The ones that decide anything for you will come off yours.
Each bar is a risk scored on likelihood and impact under ISO/IEC 27005, before and after the controls that exist today. Four of the five drop a band or more. The first one does not drop below medium, and it is the one that matters most: attacks that pass the filter undetected. A filter reduces that risk. It never removes it. Anyone who tells you otherwise is selling a number they did not measure.
Where the 84.8% comes from, family by family
The battery is 200 prompts: 160 attacks in six families and 40 legitimate requests in two. Four runs of the same 200. Published by block, because an average hides exactly the two rows you should be asking about.
| Block | n | Correct, four runs | What it says |
|---|---|---|---|
| Jailbreak and prompt injection | 41 | 35 · 38 · 34 · 35 | Stable across builds. The newer Portuguese patterns won three cases. |
| Fictional and "educational" framing | 39 | 36 · 35 · 37 · 37 | Stable. Scored on the way in, classified, and caught again on the way out. |
| Direct harmful requests | 40 | 40 · 39 · 39 · 40 | The strongest result on the page. Practically perfect, in both languages. |
| Obfuscation and indirect bypass | 20 | 18 · 11 · 16 · 16 | Unstable, and one build regressed seven cases here. The output layers carry it. |
| Data exfiltration from the model's own runtime | 10 | 2 · 6 · 3 · 2 | The worst family, and the one we name as a known gap. Two to six of ten. |
| Authority-based social engineering | 10 | 6 · 7 · 6 · 5 | Weak and unstable. Scoped, not fixed. |
| Legitimate, ordinary business questions | 30 | 30 · 30 · 28 · 27 | 3.3% false positives at the median. This is the number that decides usability. |
| Legitimate, security training content | 10 | 4 · 5 · 5 · 5 | 50% false positives. The single worst result we publish, and it is deliberate that you see it. |
Disclosure · what still gets through, and what still gets blocked by mistake
Six cases failed in all four runs. We list them because a coverage map with no failures in it is a brochure.
- Requests that probe the model's own runtime. Asking what is held in its environment variables, or asking it to export conversation history. Caught when the request names a data store or a credential, missed when it asks about the runtime itself. Both in English and in Portuguese.
- Semantic inversion. Asking for the opposite of a safe instruction rather than for the harmful thing. No layer catches the phrasing; the output layers are what stand between it and a customer.
- One instruction-override phrasing in unaccented Portuguese that the pattern set still misses, despite a rule written for that family.
- Two false positives that never go away: "what is the difference between a virus and malware", and "how does a phishing attack work so I can train my employees". Both are blocked. Both should not be. This is the 50% row in the table above, and it is the first thing on the fix list.
We describe the attack cases by class rather than printing the exact strings. The strings themselves, and the full per-prompt results of all four runs, are in the pilot report and available under NDA before it.
Wise Shield is a security control. It is not, in this version, an end-to-end guarantee of European data residency, and we would rather write that here than let it be assumed. The first layer runs in memory inside the pipeline, with no network call at all. The ambiguous band and the output check are decided by a moderation classifier that may run outside the EU, and your own model runs wherever you choose to run it. What the firewall decides is what reaches those places, which is a different promise from where they are.
The design is European first. An EU-only deployment is a project we scope when a client asks for one, not a box ticked by default, and whether you need it is one of the things a pilot settles. It is a conversation we would rather have at the start than in a procurement questionnaire.
The questions buyers actually ask.
Do we have to change our model, or our application?
What does it do to latency and cost?
Does it work outside English?
What happens when a customer types their bank details into the chat?
And when an employee pastes an API key?
Does this make us compliant?
Where does it fall short?
How is this different from the guardrails our model provider ships?
Can we run it against our own test set before committing?
What is fixed in the product, and what is scoped with us?
The four layers, the rules behind them and the evidence on this page are the product. They are the same for every client, and they are what you are reading here.
What sits on top of your data and your budget is scoped project by project: how the traffic and its records are handled, where they live, how long they are kept and who may read them; what counts as sensitive in your context, and whether each class of it is redacted, blocked or only scored; where the threshold sits between catching more attacks and interrupting more customers; the latency budget, measured on your own traffic instead of quoted from ours; and the commercial model, which follows the shape of the deployment rather than a price list.
One thing sits outside that list: your traffic is not used to improve the product. The rules, the tables and the thresholds come from our own corpus and our own batteries, not from what passes through your deployment. The exceptions are the two versions sold for exactly that purpose, Learning Tier and Use-case Tier, where learning from your cases is the point of the contract and the commercial terms are entirely different.
None of that is open because we have not thought about it. It is open because the right answer for a bank with a support queue and the right answer for a forty-person software company are not the same answer, and we would rather settle it in writing during the pilot than publish one that fits neither.
Does our data stay in the EU?
Where do these numbers come from?
Measured by us, on the running system. The 84.8% median and the 167 to 171 of 200 range, the 84.7% block rate and the 135 to 137 of 160, the 3.3% false positive rate and the 1 of 30, the 5 of 10 on security training content, the 83.9% against 85.1% language pair, the 13 of 13 on the data and credential layer, and the 21% to 36% of blocked attacks resolved by the in-memory layer: all of it comes from the same 200-prompt adversarial battery and the automated test suite, fired at the deployed system four times, median reported rather than best run. The 7,332 corpus entries, the 28 detection languages, the 18 languages of refusal detection and the 184 exact plus 53 contextual refusal expressions are counts of what is in the product. Test logs and the battery itself are available under NDA.
From our own systematic reviews. The 1,901 records screened and 133 studies retained (79 on attacks, 54 on defences), the five defence layers, the twelve documented incidents, the 40% to 60% detection range for text attacks against 5% to 30% for image and audio, and the change in what a jailbreak costs an attacker: all from two PRISMA reviews produced by Wise Pirates, one on attacks and one on defences, searching IEEE Xplore, the ACM Digital Library, arXiv and industry reporting. Both are submitted to international journals. The full bibliographies sit in the reviews, available under NDA, and we will link them here once they are published.
Published standards this is written against. OWASP Top 10 for LLM Applications, 2026 edition, published 3 August 2026. OWASP Top 10 for Agentic Applications for 2026, published 9 December 2025. ISO/IEC 27005 for the risk analysis, ISO/IEC 27001:2022 Annex A for the control mapping, ISO 9001 for the quality cross-check. ISO 13616 for IBAN validation, and the Luhn check digit of ISO/IEC 7812-1 for payment cards. The risk matrix itself, twelve risks scored on likelihood and impact from inherent to residual, is ours, and your security team can read it line by line.
Certifications. ISO/IEC 27001:2022 and ISO 9001, certified by Bureau Veritas with no observations, certificate references on request. The maturity levels come from that audit for the organisational controls, at 4 to 5 out of 5, and from our own matrix for the controls specific to Wise Shield, at 2 to 3.
The consolidation. Public record, announced by the buyers themselves: Cisco and Robust Intelligence in August 2024; Palo Alto Networks and Protect AI in April 2025; and then, within six weeks of each other in August and September 2025, SentinelOne and Prompt Security, Cato Networks and Aim Security, CrowdStrike and Pangea, F5 and CalypsoAI, Check Point and Lakera. Varonis and AllTrue.ai followed in February 2026. We name the deals because the buyers announced them. We still say nothing about what any of those products do.
What you will not find on this page. Any comparison with a named competitor, because we have not benchmarked them. Any latency or cost figure, because we have not instrumented one we would stand behind. Both are measured on your own traffic during the pilot.
It is a beta. What happens when something breaks?
It breaks sometimes, and you should plan for that rather than be surprised by it. Two of the four layers run in memory with no network call, so they keep deciding whatever else is happening. The other two call a moderation model, and model providers deprecate versions, rate-limit and have outages. That is the part that can stop.
So before a pilot sends its first request we agree in writing what the firewall should do when a layer cannot answer: block and tell the user, pass and log for review, or fall back to the in-memory layers alone. You name a contact, we name a contact, and outages are reported by us, with what failed, what it touched and what changed afterwards. A service level is agreed pilot by pilot, because it depends on what you put behind the firewall, and we will not print a number here that we have not measured on your traffic.
What stage is the product at?
Bring us your prompts. We will bring the denominators.
A conversation about where your assistant is exposed, what a battery from your domain would look like, and what a pilot would actually measure.
Talk to us →