AI agent outcome verification: prove what your agents actually did.
Your agent says it completed the action. Pruvz checks the systems of record and tells you whether the business result really happened: each high-impact action is captured with the exact policy in force at decision time, independently read back from the systems of record, classified, and routed to human review when the evidence disagrees.
A successful tool call proves the agent invoked an API and the API answered according to its contract. It does not prove the refund settled, the account changed, or the claim closed: downstream systems hold, reject, transform, and partially process requests after reporting success. When an agent's work has financial, customer, or compliance consequences, someone eventually asks for proof that is independent of the agent's own report. Outcome verification is that proof, and producing it is the one job Pruvz exists to do.
How verification works in Pruvz.
Five steps, each visible in the recorded end-to-end product demo and each leaving an entry on an ordered, append-only evidence trail.
The claim is captured
When an agent takes a high-impact action, Pruvz records what the agent claims it did, together with the decision-time context: what the agent saw, what it decided, and the action it intended to execute.
The policy is snapshotted
The exact policy in force at decision time is captured and embedded with the action. Six months later the question is never what the policy says today; it is what applied at that moment.
The outcome is read back, independently
After execution, a separate verification worker reads the systems of record directly, read-only, inside a defined verification window with retries. The agent's own report never decides the result.
The outcome is classified
Each action ends in exactly one terminal business fact: verified, outcome mismatch, or verification failed. A source that cannot be read yet keeps the action verification-pending; it is never misclassified as a business mismatch.
Mismatches go to human review
An outcome mismatch enters a review queue with the complete evidence attached. The reviewer's decision is appended to the record as a new fact; it never rewrites the verification result.
A window with retries, then a final answer.
Verification runs inside a defined window: a verification job is created when execution completes, and the worker reads the systems of record with retry and backoff until it observes a terminal state. The result becomes final on a terminal observation, and a final result states what the systems of record showed during that window. Pruvz does not silently re-check outcomes afterwards and does not claim continuous monitoring: later facts, such as a human review decision, are appended to the evidence as new entries, and automatic re-verification after a source-system change is a planned capability on the roadmap, not current behavior.
The verification lifecycle, documented on the security pageWhat the record looks like.
Every verified action produces an evidence packet: the agent's claim, the decision-time policy snapshot, the execution receipt, each independent system-of-record read-back with its server-assigned trust level, and the derived verification result, in an ordered, append-only sequence. A real packet, captured from the product demo, is published and explained field by field, with a public schema and validator you can run against it on your own machine. Verified outcomes then roll up into a business view, with drill-down from any number to the evidence behind it.
Scope, stated plainly.
In the current product demo, read-back runs against a billing system and a CRM system, and the demo page documents exactly how. Production connectors to live systems of record, such as Stripe and HubSpot, are the scope of the founding design-partner program and will be selected and developed with the founding design partners; the verification engine is designed for systems of record such as CRM, billing, ERP, and ticketing. Cryptographic sealing of exported packets is planned and is not claimed as a current fact: the current integrity guarantee is the ordered, append-only evidence trail. And Pruvz never approves, blocks, or corrects agent actions: it is the record, not the enforcement.
How this compares to the other ways teams verify outcomesCommon questions.
What exactly does Pruvz compare when it verifies an outcome?
Pruvz compares what the agent claimed against what the systems of record actually show. For each high-impact action it captures the agent's claim and the decision-time policy, then independently reads the relevant systems of record after execution and checks the material fields of the outcome. Only that independent read-back assigns the final verification result; claims made by the agent or by external callers are recorded as claims and can never be counted as independent evidence.
Does Pruvz sit in the agent's execution path?
No. Pruvz is non-blocking by design: it records and verifies off the agent's critical path, and it does not approve, block, or slow down agent actions at runtime. It also does not execute corrections in external systems. When verification finds a mismatch, the case is routed to human review and resolved through your normal business channels, with the resolution appended to the evidence.
What happens when an outcome cannot be verified?
The action stays verification-pending while the verification window is open: unreadable or non-terminal observations are never classified as business mismatches. If retries are exhausted on a source that cannot be read, the action is classified as verification failed, which is a technical outcome, distinct from an outcome mismatch. Both states are visible, so silence never reads as success.
Is AI agent outcome verification available in Pruvz today?
Pruvz is a working product: the full flow, from agent action to independent system-of-record read-back, outcome classification, and human review of a caught mismatch, runs end to end in the product demo against a billing system and a CRM system. The founding design-partner program is open: production connectors to live systems of record and the initial partner workflows will be developed with the founding design partners.
See a verified outcome, and a caught mismatch.
The fastest way to evaluate outcome verification is to watch it disagree with an agent. In the recorded demo, one run verifies cleanly and one run catches an agent that reported success when the refund never happened. The founding design-partner program is open for 3-5 teams running agents with real business impact.