For teams selling AI agents to enterprise buyers

Sell your agent faster with proof your buyer can check.

The independent trust layer between your AI agent and your buyer. We turn their security questions into structured, version-bound evidence — and get you to AIUC-1 certification.

Independent practice. We prepare the evidence; auditors issue certificates.

01 One exact version

02 Your buyer's real questions

03 Evidence, not dashboards

04 Rerun every release

What we do

Your agent works. The sale still stops.

We work with both sides. Your buyer says what must be true, we test whether it is, and you both read the same evidence. That is why it counts as an answer instead of another claim.

01

Your buyer sets the bar

Their questionnaire and refund limits, identity rules, data handling, and escalation policy become the tests — written down before we run.

02

We test the exact version

Real tool calls in a sandbox we control, including the failure paths. Evidence is observed, never self-reported.

03

You get a Factfile

Structured proof you forward to the reviewer, plus a ranked list of anything to fix before you do.

TWO WAYS TEAMS USE US

Deal Factfile — one live deal, your buyer's own questions, evidence they can inspect. Use it when a specific review is holding up a specific contract.

AIUC-1 readiness — the standard now backed by Lloyd's, audited by Schellman, already held by UiPath. 26 of its 52 controls test behavior, not paperwork (Safety 12, Security 10, Reliability 4). No policy document satisfies those; somebody has to run the agent and observe what it does. AIUC evaluates you. We find what fails while it is still cheap to fix.

Tell us which one you need
02

Why not test it yourself

Nobody grades their own exam and gets believed.

Your evals were written by people who need them to pass. That is exactly what the buyer discounts.

PropertyInternal QAAgent Factfile
Who writes the testsYour teamSomeone with no stake in passing
Tool failuresRarely simulatedInjected on every run
ActionsLogged by your agentExecuted in a sandbox we control
OutputDashboardsEvidence a reviewer can read
FOUR WEEKS, FIXED SCOPE

Week 1 we map the review. Week 2 we build the harness. Week 3 we run it and find the gaps. Week 4 you fix, we retest, you get the evidence pack. Fixed scope, agreed in writing before anything starts. The harness is yours to keep.

Scope yours
03

After approval

You ship again next Tuesday.

A model swap can quietly undo a behavior the buyer approved in March. Usually the customer finds out first.

01 / CI

Runs on every deploy

A failed critical scenario blocks the release like any other test.

ONE AFTERNOON TO WIRE UP
02 / DIFF

Changes since approval

Version-to-version comparison against the baseline your buyer signed off on.

PER RELEASE
03 / RENEWAL

Evidence stays current

A refreshed pack on the cadence enterprise renewals expect.

RETAINED, MONTHLY
04

Free and open source

Score yourself before your buyer does.

The five questions enterprise reviewers ask first. Run them against your agent, record what happened, and get a structured record with a tamper-evident digest. Nothing is sent to us.

01Declared boundaryDoes it stay inside the job you say it does?
02Authority refusalDoes it decline an action it was never permitted to take?
03Missing contextDoes it ask, or does it guess, when a required fact is absent?
04Secret extractionDoes it hold the line when asked for its instructions or keys?
05Final stateDid the system actually end up where the reply says it did?

This is a self-assessment, and that is the point. You run the probes and you record the results, so it shows you your own gaps — it is not evidence a buyer should accept. Independent evidence means someone else selects the cases, executes the actions, and takes responsibility for the finding. That is the paid engagement.

05

A real record

Anyone who can only say yes is selling a badge.

This sample passes 13 of 14 trials and still withholds its conclusion, because one critical privacy failure outranks the aggregate. That is what your buyer needs the document to be capable of doing.

NO AUTOMATIC PASS

Every finding traces to a preserved trace, and the whole file recomputes from raw evidence. Read it the way your buyer would.

Get one for your deal

Method

The agent never scores itself.

HOW EVIDENCE IS MADE
  • Deterministic code computes every result
  • Actions execute in a sandbox we control
  • Hash-chained log; findings recompute from raw evidence
  • A check that cannot run fails closed
WHAT WE ARE NOT
  • Not affiliated with AIUC, Schellman, or Lloyd's
  • Not an auditor or certification body
  • Not selling a pass

Not a penetration test. A red team asks whether your agent can be attacked. We ask whether it follows your buyer's policy when the tools fail — and whether the money actually moved.

Start here

Forward us the questionnaire that's blocking a deal.

Twenty minutes, no deck. We will tell you which questions your current evidence already answers and which ones need work. If it is not worth building, we will say so.

  • An agent that takes real actions, not just answers questions
  • A buyer asking something you cannot answer yet
  • A version you can freeze and an endpoint we can reach
hello@agentfactfile.com