If you cannot observe it, you cannot operate it.
Independent · AI-native · ISO 9001 + ISO 27001
One obsession: knowing what your AI agents did, why they did it and what it cost, before a customer tells you. An agent that works but cannot be seen is not in production. It is unsupervised. We make autonomous agents observable the way we run our own: every run traced, every run evaluated, every silence questioned.
What agent observability means, in one minute.
The rest of this page goes deep, and it is written for engineering, data and security teams. If you read only one part, read this: we trace every run, evaluate every outcome, and question every silence, so you can answer "why" before a customer does.
See every run, end to end
A span for every model call, tool call and retrieval, on the OpenTelemetry GenAI conventions, replayable months later.
Judge the outcome, not the status
Success is evaluated, by rules where rules exist and calibrated judges where they do not, never read off a green light.
The run that did nothing
A clean exit that wrote nothing is a state, not a quiet day. We baseline normal, so silence becomes a signal.
In one line: we make autonomous agents observable the way we run our own, with a senior human who can answer "why" when it counts.
Not application monitoring with a new dashboard.
Your monitoring stack was built for requests: they return in milliseconds, they succeed or they fail, and they do the same thing every time. An agent does none of that. It runs for minutes, decides its own next step, and can return a perfect-looking answer that is wrong.
The unit of work is a run, not a request
One run is model calls, retrievals, tool calls and retries, sometimes over hours. Uptime and error rates describe the servers underneath. They say nothing about whether the agent did the right thing or stayed inside its budget.
Success is a judgement, not a status code
In agentic systems the most expensive failures arrive with a green light, so correctness has to be evaluated, by rules where rules exist and calibrated judges where they do not, never read off a status.
Non-deterministic by nature: the same input can take a different path twice, so the same bug produces a different trace every time, and grep does not find it.
The first incident will not be an outage. It will be a quiet month.
Autonomy changed the shape of the risk. Nothing goes down. The agent keeps working, keeps spending, and keeps being slightly wrong, until someone outside the company notices. The market has adopted tracing, but it has not yet connected the traces to a judgement of quality.
of agentic AI projects expected to be cancelled by end 2027, for cost, unclear value and weak risk controls, all observability problems.Gartner
run online evaluations on live traffic, where the real failures actually live.LangChain, 2025
wrong assumption repeated thousands of times overnight is not a bug, it is a month of cleanup.
tokens are a budget line now. A budget that only reports at month end is a receipt, not a control.
Not only crashes. Also the answer that looked fine.
A crash is the easy case: it pages someone. The failures that cost money in agentic systems sit at the quiet end, and they share one property: every conventional signal stays green while they happen.
A span for every decision, from the trigger to the action.
A trace is not a log line that says "agent finished". It is the ordered tree of everything the agent read, decided and did, with identity and cost on every node. This is the record we build for every run, on the OpenTelemetry GenAI conventions.
This wraps the agents your Process Automation & Agentic AI team builds: vendor-neutral spans that land in the backend you already run, reconstructable months later. The traces you run on are the evidence you will be asked for.
Eight signals that say an agent is going wrong. And the ninth that says nothing happened.
These are the measurements we instrument on every agent we observe. Together they answer the four questions a board, an auditor and an engineer all ask: what did it do, was it right, what did it cost, and what changed.
Trace completeness
Every run reconstructable end to end. A gap in the trace is itself a finding, because it is where the next incident hides.
Task success, evaluated
Not assumed from a status code. Rules where rules exist, a calibrated judge where they do not, humans on the sample that matters.
Tool-call health
Error, retry and latency for every tool and MCP server. The tool that quietly returns nothing is more dangerous than the one that throws.
Steps and loops
Step count and repetition against the baseline, so a re-planning loop is cut on its count long before it is cut on its cost.
Cost per task
Tokens rolled up to the outcome, not the token, so leadership sees cost per resolved case and engineering sees which step doubled it.
Latency and duration
Duration monitored in both directions, because a run that finishes far too fast is as suspicious as one that hangs.
Drift and change
Model version, prompt hash, tool schema, retrieval index, input distribution, each a visible event, so "what changed" already has an answer.
Guardrail & escalation
Policy hits, tier changes and human escalations, so the moments an agent was stopped or handed over are themselves measured.
Silence
Every agent has an expected volume. A run that exits clean and writes nothing, makes zero tool calls, or produces nothing, is a state, not a quiet day.
Cost per task, not per token. And a budget that acts.
A token bill tells you what the model charged. It does not tell you what a resolved ticket costs, which agent is burning the budget, or that one run at three in the morning spent what the whole day should have. So we instrument cost like the behaviour signal it is: a daily budget per agent, a warning at 80%, and an automatic pause at 100%, with a cap on steps so a loop is cut on its count first.
Evaluation is the engine, not the gate at the end. A golden set labelled by your domain experts runs as a regression before any change ships, and a calibrated judge scores live traffic continuously, so the long-tail failures surface where they actually live.
Security sees the threat. Observability sees the behaviour.
Autonomy is only worth deploying when you can see it run. So observability is not a separate tool bolted next to security, it is the same system seen from another angle, built together by the same team. One protects the agent from what is done to it; the other shows what the agent itself is doing.
The traces feed the guardrails
What the agent actually did, span by span, is exactly what the guardrails and policies need to decide what it should be allowed to do next.
The baselines feed anomaly detection
The baseline that turns forty steps into a signal is the same baseline that flags a compromised or misbehaving agent before it does damage.
The kill switch is shared
Circuit breakers by cost and by count are wired to the same kill switch security uses, so a run that goes wrong is stopped once, cleanly.
The evidence is one record
The traces engineers debug on are the logs an auditor asks for. Observability built for one is evidence built for the other, designed once.
See Agentic Security → the controls the traces feed, and the kill switch they pull.
Three silent failures we caught in our own pipelines.
We run autonomous engines every morning: ingestion, scoring, reporting, review. Each of these happened to us in 2026, each with a green status, and each changed how we build. A page about observability that only lists other people's failures would not be observable itself.
The run that finished in 29 seconds
- It finished in under half a minute with exit code zero, having lost access to its credentials at startup, and wrote a full day of empty values.
- A downstream rule read the emptiness as an improvement and recorded a win. Sibling jobs wrote the error inside their output and exited clean, once an hour, for fifteen hours.
What changed: abort before writing when a secret cannot be read, treat an all-empty result as an outage, and monitor duration in both directions.
The step that timed out and was marked complete
- The step timed out, the error was caught and logged inside the output, and the orchestrator marked the task complete and moved on.
- The run looked successful end to end while a chunk of the work had quietly failed, invisible to every status-based check.
What changed: a caught error is still a failure, evaluated on the outcome, not the exit code, and every step reconciles what it produced against what it should have.
We run agents every day, so we know how they fail quietly. The lessons on this page are ours, and the controls we sell are the ones we needed first. A third lesson, on a right answer that cost thirty times what it should, is walked through in the deep dive.
Six parts. Skip to the one you need.
Everything here is written for engineering, data and security teams, and each part stands on its own.
Tracing and replay
OpenTelemetry GenAI spans on every invoke, model call, tool call and retrieval, stored and replayable, redacted and retained by policy.
Evaluation in production
Offline golden sets before a change ships, and online evaluation on live traffic with a judge calibrated against human labels.
Cost and token telemetry
Cost rolled up to the outcome, budgets that warn and pause, and loops cut on their count before their cost.
Baselines, drift and silence
A baseline per agent and task that turns telemetry into alerts, with drift and silence as first-class states.
Alerting and the human loop
Three tiers, an automatic first response, and every alert carrying its runbook, so who sees the event is who acts on it.
Continuous improvement
Failures clustered by meaning, every incident becomes a test, and each fix goes back into the golden set and the guardrails.
The traces you run on are the evidence you will be asked for.
Observability built for engineers and evidence built for regulators are the same records, designed once. For high-risk systems the EU AI Act makes automatic logging a design requirement (Article 12), covering risk-relevant events and post-market monitoring, and deployers keep those logs for at least six months (Article 26). We design traces to be that evidence: complete, time-stamped, attributable to an agent and a version, retained under a written policy.
Traces on the vendor-neutral OpenTelemetry GenAI conventions, measurement mapped to the NIST AI RMF, and monitoring mapped to ISO/IEC 42001. One caveat, stated rather than hidden: the OTel GenAI conventions are still in development status, so we pin versions and track the spec as it settles.
Where our approach is different.
Everyone has logs. Few can say whether the agent was right. This is the gap we are built for.
| Capability | APM / monitoring | Generic tracing tools | Wise Pirates |
|---|---|---|---|
| Run-level, not request-level | Request-level | Run-level | Run-level, evaluated |
| Success evaluated, not read off a status | Status code | Rare | Owned, judged |
| Silence and empty-result detection | No | Rare | Owned |
| Cost per task, with budgets that act | No | Cost dashboards | Owned, enforced |
| EU AI Act evidence, designed in | No | No | Built in |
| One system with Agentic Security | No | No | Shared kill switch |
| Run by people who run agents daily | A vendor | A tool | Operators, evidence-first |
Tools give you tracing. We connect the traces to a judgement of quality, a cost you can control, and a human who can answer why. Observability that is also evidence, and that shares a kill switch with your security.
A partner who runs agents, not just tools for them.
Observability by default: one of the four principles every Wise Pirates engagement is built on, in the first sprint, not the retrofit after the first incident. Evidence, not opinions: OpenTelemetry GenAI conventions, the NIST AI RMF, ISO/IEC 42001 and the EU AI Act, so what you see is also what you can show. Every figure on this page has a named source, and every alert we raise has a run behind it.
We run agents every day, so we know how they fail quietly.
Operators, not just builders
We design agents for clients and run autonomous pipelines of our own, daily. The controls we sell are the ones we needed first.
Evidence, not opinions
OTel GenAI, NIST AI RMF, ISO/IEC 42001 and the EU AI Act. We would rather show you the trace than ask you to trust the summary.
One system with security
The traces feed the guardrails, the baselines feed the anomaly detection, and the kill switch is shared. Built together by the same team.
Autonomy is only worth deploying when you can see it run. That is the whole game.
Observability connects to the rest of the agency.
Autonomy is only worth deploying when you can see it run, so observability sits next to everything agentic we build and secure.
Built in the language of observability.
Depth here is not only strategy, it is vocabulary. We operate inside the terms your engineering, data and security teams use every day, and that fluency is how you earn credibility with a technical audience.
Spans, evals, baselines, cost and runbooks, in the language your teams already speak, applied to agents you can actually operate.
The questions we hear most.
What is agent observability, and how is it different from monitoring?
Why is this a board-level risk, not just an engineering concern?
Everyone has logs and tracing now. Is that enough?
What signals do you actually instrument?
How does observability connect to Agentic Security?
Do you help us meet the EU AI Act and standards?
Do you build the agents, or only observe them?
How do we start?
Give your agents autonomy you can actually see.
Start with a free agent observability assessment. Tell us what your agents do today and what you can see of it, and a senior engineer maps the blind spots and where to start, at no cost. It is the same assessment we open a paid engagement with.
Request my free assessment →See Agentic Security →