We test whether your agent gives the right answer — and we report a failure only after it has reproduced three times out of three, for the same reason each time.
Your agent answers questions about your policies all day — return windows, eligibility, fees, cover. You find out it got one wrong when a customer tells you, or when someone acts on the answer.
Nobody can prove it is right, because the thing that would prove it — a transcript, a run hash, a reproduction log, produced by someone with no stake in the result — does not exist until somebody makes it. That is what we sell.
Findings we could not reproduce are named as unreproduced and excluded from the count. They are still handed over — we just do not let you act on them.
Rather than describe it: read a complete audit report, in the form a client receives it.
Standard RAG metrics reward an answer that faithfully quotes a stale clause. So we do not grade against the retrieved text — we grade against a canonical file that decides what is true before the agent ever runs. Traps go in both directions: sometimes the outdated number is the generous one.
Conflicting policies. A clause superseded last quarter. A question in Russian against a rule written in English. Boundary values one day either side of a threshold. And a user trying to talk the agent into a discount it is not allowed to give.
A grader that looks correct fails in both directions. Ours once marked a textbook injection refusal as a failure, because the pattern matched the agent quoting the payload it was rejecting. Every failure is read by hand before it is written down. Where a rule needs judgement, the model judging it is calibrated against human labels first.
Each failing case is run three times. If it does not fail all three times for the same reason, it does not become a finding. On our own reference audit this gate removed 24 of 29 candidate failures — a report without it would have been mostly noise.
Reproduced 3 / 3We ran the method against a support agent we built ourselves, published the findings, the limitations and the code, and left the corrections visible. You can check the work before you hire us.
The negative result matters as much as the findings: our own taxonomy predicted that multi-turn conversations degrade accuracy. We measured it and found no degradation. We published that too, and corrected the taxonomy rather than the data. Read the full report and finding register.
Between March 2025 and May 2026 the AI evaluation tooling market consolidated into the hands of the vendors being evaluated. OpenAI acquired Promptfoo. Anthropic absorbed Humanloop and the product was wound down. Cisco is folding Galileo into Splunk.
An audit is worth what its independence is worth, so ours is stated plainly and can be checked:
The full statement — funding, suppliers, and what would have to change for it to be wrong: independence.
50% on signature. Re-testing after your fix is a separate line, never bundled into the first price. Introductory pricing while we build our first public references.
You pay half and keep the corpus. A clean result is a real result — but note that this term gives us an incentive to find failures. We know that, and we are not hiding it. What holds us back is the three-out-of-three gate, and the code behind that gate is public. Audit our filter before you trust our findings.
Companies running a customer-facing AI agent in a regulated market — a challenger bank, an insurer, a telecoms or energy provider, a pension or health service. Your agent applies your rules to your customers, and a wrong answer is your cost, not your supplier's.
The usual trigger is one of four: a customer acted on an answer that was wrong, an internal review asked how you know the agent is correct, the EU AI Act obligations that took effect in August 2026 landed on your team, or you are about to widen what the agent is allowed to decide.
We are not the right call if your agent is still a pilot with no customers on it. Come back when it is answering real questions.
Twenty minutes. Tell us what broke and how you found out. If an audit is not the right thing for you, we will say so — and we do not do the implementation work, so we have no reason to talk you into one.