Independent audits of AI agents.

We test whether your agent gives the right answer — and we report a failure only after it has reproduced three times out of three, for the same reason each time.

Your agent answers questions about your policies all day — return windows, eligibility, fees, cover. You find out it got one wrong when a customer tells you, or when someone acts on the answer.

Nobody can prove it is right, because the thing that would prove it — a transcript, a run hash, a reproduction log, produced by someone with no stake in the result — does not exist until somebody makes it. That is what we sell.

The deliverable

What you get

Findings we could not reproduce are named as unreproduced and excluded from the count. They are still handed over — we just do not let you act on them.

Rather than describe it: read a complete audit report, in the form a client receives it.

Method

Four steps, one gate

01

We build ground truth you cannot argue with

Standard RAG metrics reward an answer that faithfully quotes a stale clause. So we do not grade against the retrieved text — we grade against a canonical file that decides what is true before the agent ever runs. Traps go in both directions: sometimes the outdated number is the generous one.

02

We ask the questions your customers will

Conflicting policies. A clause superseded last quarter. A question in Russian against a rule written in English. Boundary values one day either side of a threshold. And a user trying to talk the agent into a discount it is not allowed to give.

03

We audit our own grader before we trust it

A grader that looks correct fails in both directions. Ours once marked a textbook injection refusal as a failure, because the pattern matched the agent quoting the payload it was rejecting. Every failure is read by hand before it is written down. Where a rule needs judgement, the model judging it is calibrated against human labels first.

04

Nothing enters the report until it repeats

Each failing case is run three times. If it does not fail all three times for the same reason, it does not become a finding. On our own reference audit this gate removed 24 of 29 candidate failures — a report without it would have been mostly noise.

Reproduced 3 / 3
Evidence

The reference audit is public

We ran the method against a support agent we built ourselves, published the findings, the limitations and the code, and left the corrections visible. You can check the work before you hire us.

185
graded cases across two corpora
29 → 5
candidate failures that survived the reproduction gate
96.7%
judge agreement with human labels · Cohen's κ 0.95
44
limitations documented, each with its direction of bias
25
contract checks every adapter must pass, including the one we write for you
4
published findings — and one published negative result

The negative result matters as much as the findings: our own taxonomy predicted that multi-turn conversations degrade accuracy. We measured it and found no degradation. We published that too, and corrected the taxonomy rather than the data. Read the full report and finding register.

Independence

Who we are not owned by

Between March 2025 and May 2026 the AI evaluation tooling market consolidated into the hands of the vendors being evaluated. OpenAI acquired Promptfoo. Anthropic absorbed Humanloop and the product was wound down. Cisco is folding Galileo into Splunk.

An audit is worth what its independence is worth, so ours is stated plainly and can be checked:

The full statement — funding, suppliers, and what would have to change for it to be wrong: independence.

Pricing

Fixed price, stated up front

Reproduction audit · 14 days $4,500 One system, one version. Every reproduced failure with its artifacts, plus the remediation list.
Full audit · 4–6 weeks from $14,000 Corpus built for your domain, re-test after your fix, and a written attestation of what was measured.
Regression watch · monthly $2,500 The same dataset on every release, with a report on what changed since the last one.

50% on signature. Re-testing after your fix is a separate line, never bundled into the first price. Introductory pricing while we build our first public references.

If we find nothing

You pay half and keep the corpus. A clean result is a real result — but note that this term gives us an incentive to find failures. We know that, and we are not hiding it. What holds us back is the three-out-of-three gate, and the code behind that gate is public. Audit our filter before you trust our findings.

Scope

What we do not do

Fit

Who this is for

Companies running a customer-facing AI agent in a regulated market — a challenger bank, an insurer, a telecoms or energy provider, a pension or health service. Your agent applies your rules to your customers, and a wrong answer is your cost, not your supplier's.

The usual trigger is one of four: a customer acted on an answer that was wrong, an internal review asked how you know the agent is correct, the EU AI Act obligations that took effect in August 2026 landed on your team, or you are about to widen what the agent is allowed to decide.

We are not the right call if your agent is still a pilot with no customers on it. Come back when it is answering real questions.

Start with a conversation, not a contract

Twenty minutes. Tell us what broke and how you found out. If an audit is not the right thing for you, we will say so — and we do not do the implementation work, so we have no reason to talk you into one.

Email
yusifli.pervin@gmail.com
LinkedIn
linkedin.com/in/yusifoff
Code
github.com/pervincinal/agentproof