Your buyer sets the bar
Their questionnaire and refund limits, identity rules, data handling, and escalation policy become the tests — written down before we run.
For teams selling AI agents to enterprise buyers
The independent trust layer between your AI agent and your buyer. We turn their security questions into structured, version-bound evidence — and get you to AIUC-1 certification.
Independent practice. We prepare the evidence; auditors issue certificates.
01 One exact version
02 Your buyer's real questions
03 Evidence, not dashboards
04 Rerun every release
What we do
We work with both sides. Your buyer says what must be true, we test whether it is, and you both read the same evidence. That is why it counts as an answer instead of another claim.
Their questionnaire and refund limits, identity rules, data handling, and escalation policy become the tests — written down before we run.
Real tool calls in a sandbox we control, including the failure paths. Evidence is observed, never self-reported.
Structured proof you forward to the reviewer, plus a ranked list of anything to fix before you do.
Deal Factfile — one live deal, your buyer's own questions, evidence they can inspect. Use it when a specific review is holding up a specific contract.
AIUC-1 readiness — the standard now backed by Lloyd's, audited by Schellman, already held by UiPath. 26 of its 52 controls test behavior, not paperwork (Safety 12, Security 10, Reliability 4). No policy document satisfies those; somebody has to run the agent and observe what it does. AIUC evaluates you. We find what fails while it is still cheap to fix.
Why not test it yourself
Your evals were written by people who need them to pass. That is exactly what the buyer discounts.
Week 1 we map the review. Week 2 we build the harness. Week 3 we run it and find the gaps. Week 4 you fix, we retest, you get the evidence pack. Fixed scope, agreed in writing before anything starts. The harness is yours to keep.
Scope yoursAfter approval
A model swap can quietly undo a behavior the buyer approved in March. Usually the customer finds out first.
A failed critical scenario blocks the release like any other test.
ONE AFTERNOON TO WIRE UPVersion-to-version comparison against the baseline your buyer signed off on.
PER RELEASEA refreshed pack on the cadence enterprise renewals expect.
RETAINED, MONTHLYFree and open source
The five questions enterprise reviewers ask first. Run them against your agent, record what happened, and get a structured record with a tamper-evident digest. Nothing is sent to us.
This is a self-assessment, and that is the point. You run the probes and you record the results, so it shows you your own gaps — it is not evidence a buyer should accept. Independent evidence means someone else selects the cases, executes the actions, and takes responsibility for the finding. That is the paid engagement.
A real record
This sample passes 13 of 14 trials and still withholds its conclusion, because one critical privacy failure outranks the aggregate. That is what your buyer needs the document to be capable of doing.
Every finding traces to a preserved trace, and the whole file recomputes from raw evidence. Read it the way your buyer would.
Get one for your dealMethod
Not a penetration test. A red team asks whether your agent can be attacked. We ask whether it follows your buyer's policy when the tools fail — and whether the money actually moved.
Start here
Twenty minutes, no deck. We will tell you which questions your current evidence already answers and which ones need work. If it is not worth building, we will say so.