An approval is a control operation, not a workflow step
Every major control framework treats authorization and approval as control activities. The COSO Internal Control framework names authorizations and approvals among the actions that carry management's directives into practice, and the U.S. Government Accountability Office's Green Book, which adapts COSO for federal entities, expects transactions to be authorized and executed only by persons acting within the scope of their authority. An approval, in this sense, is a preventive control: someone with the right authority confirms that a transaction is valid and within limits before it takes effect.
Controls, in turn, are judged on evidence. Section 404 of the Sarbanes-Oxley Act requires management of public companies to assess the effectiveness of internal control over financial reporting every year, and for many filers an independent auditor attests to that assessment. The auditor's standard for that work, PCAOB AS 2201, tests whether each selected control operated as designed, ranks the available procedures from the least evidence to the most (inquiry, observation, inspection of relevant documentation, re-performance), and is explicit that inquiry alone does not provide sufficient evidence that a control was effective. Asking the approver is never enough; the auditor inspects what the approval left behind.
A human approver who kept thin records can at least sit for that inquiry, however little weight it carries. An AI agent cannot be interviewed about a decision it made last quarter: there is no recollection to probe and no judgment to have explained after the fact. Auditors will still walk through the control's design and test the IT general controls around the agent, and they can re-perform an approval, but only from the inputs and rules that were captured when it happened.
Nor does the usual shortcut for automated controls carry over cleanly. AS 2201 lets auditors carry forward their testing of an entirely automated control while the IT general controls hold and the control is unchanged, because such controls are generally not subject to breakdowns due to human failure. An agent whose output can vary on the same input fits that premise poorly. Whatever record the agent's approval produced at the moment it approved therefore carries most of the evidence about each decision. That inversion is the point of this guide. When an agent grants approvals, evidence stops being the paperwork around the control and becomes the control's primary witness.
The checkbox problem arrives at machine speed
Control testing already has a name for approval evidence that proves nothing. A recurring finding in PCAOB inspection observations is insufficient testing of controls that include a review element, and specifically of whether the review operated at a level of precision sufficient to prevent or detect material misstatements. Behind that auditor-facing finding sits a company-facing gap: a signature or an "approved" status proves that someone clicked, not that anyone examined the transaction against a standard. A reviewer who cannot show what they looked at, what threshold they applied, and what would have made them say no has evidence of activity, not of control.
An agent in the approval seat multiplies that gap by volume. A person rubber-stamping approvals produces dozens of empty records a day; an agent produces thousands, each one a boolean in a database, and the aggregate looks like a control that operated flawlessly right up until someone asks what the control actually checked. The precision question does not soften because the reviewer is software. It sharpens, because more of the answer can be captured than ever could be for a person, though not all of it in the same way. The inputs the agent received and the rules in force are machine state that either was captured at decision time or was lost. The result of each check is machine state only when the check actually ran as code or as a tool call, such as a limit comparison or a three-way match. The explanation a language model writes about its own decision is something else: the agent's account of itself, useful context but not proof that any rule was applied. A defensible record keeps the two apart, measured check results as evidence and the agent's stated reasoning as context. The instrument for capturing the standard the agent applied is a decision-time policy snapshot: the approval limits, matching rules, and instructions in force at the moment of the decision, preserved with the record rather than re-derived later from whatever the policy says now.
Four questions every agent approval record must answer
Strip away the framework language and an approval defends itself months later by answering four questions. Each maps to a concrete element the record either contains or does not.
| Question | What the record must show |
|---|---|
| Authority: was the approver allowed to approve this? | The agent's identity, the authority granted to it (action types, amount limits, scope), and that this approval fell inside that grant |
| Basis: what did the approver examine? | The decision-time context the agent evaluated: the fields of the invoice, expense, discount request, or credit memo that the decision turned on |
| Standard: what rules were in force? | The policy snapshot at decision time: thresholds, matching requirements, exception criteria, and the results of the checks that actually ran, not the policy as it reads today |
| Outcome: did the approved thing happen, as approved? | The executed action, and the result confirmed independently in the system of record, matched against what was approved |
The authority question deserves more care with agents than teams tend to give it. Human approvals inherit authority from a delegation-of-authority matrix: a named person holds a named limit, and the org chart backs it. An agent's authority is whatever its credentials and instructions let it do, which in loosely configured deployments is far more than anyone consciously delegated. The record has to show the grant, not just the act: this agent, authorized for this class of approval up to this limit, approved this item within that boundary. An approval the record cannot place inside an explicit grant is not evidence of a control; it is evidence that actions were taken without one.
Segregation of duties does not dissolve because the duties are automated
The oldest rule in internal control is that the person who authorizes a transaction should not be the person who executes it or the person who records it. The Green Book frames it as separating authority, custody, and accounting so that no one individual controls all key aspects of a transaction. Automating the duties does not retire the rule; it just makes the collapse quieter. When one agent process, running one set of credentials, drafts the purchase requisition, approves it, executes the payment, and writes the log entry that says everything went well, the organization has rebuilt the exact concentration the rule exists to prevent, at higher speed and without the social friction that sometimes catches a human doing all four jobs.
Two separations carry most of the weight. The first is between requesting and approving: an agent that originates a transaction should not be the authority that approves it, any more than an employee should approve their own expense report. Distinct agents, distinct identities, distinct credentials, so the record can show that the separation held. The second is between acting and recording, and it is the one agent deployments miss most often. If the only account of an approval is the approver's own log, the approver is attesting to itself. PCAOB AS 1105, the audit evidence standard, draws the same kind of distinction at the level of the company: in general, evidence from a knowledgeable source independent of the company is more reliable than evidence obtained only from internal company sources, and the reliability of information the company generates itself increases when the company's controls over that information are effective. Applied one level down, the same logic favors a record kept by a layer that is independent of the agent, ordered and protected against silent revision, and populated with outcomes read from the systems of record rather than from the agent's self-report. That separation is an architectural advantage, not an automatic grant of audit weight: the record is still built from the company's own systems and data, so under the same standard the auditor is expected to test its accuracy and completeness, or the controls over it, before using it as evidence.
Approval is not execution: the gap after the sign-off
Approval evidence traditionally ends at the moment of sign-off, because execution used to be watched by the same people who approved. With agents executing, the interesting failures live after the approval. The approved amount and the executed amount can diverge: a $500 credit memo approved, $5,000 posted, through a unit bug or a mangled field. A retry can execute an approved action twice. An approved payment can silently never happen, which surfaces weeks later as a vendor escalation. An action can execute against the wrong entity while reporting success. In each case the approval record is pristine and the control still failed, because the thing that occurred is not the thing that was approved.
Closing that gap means extending the record past the sign-off: bind the approval, the executed action, and the outcome into one ordered record, and confirm the outcome by reading it back independently from the system of record rather than accepting the executing agent's success response. The comparison is mechanical: the amount, currency, entity, and status the system of record shows, against the amount, currency, entity, and status that were approved. A match makes the approval defensible end to end. A mismatch is precisely the case a reviewer should see, with the full chain attached.
The auditability question is already being asked
None of this is hypothetical scrutiny. In July 2024, PCAOB staff published observations from their outreach on generative AI in audits and financial reporting. Some of the audit firms told the staff that their policies for supervising and reviewing audits had not changed with the technology: an engagement team member who uses a generative AI tool is still responsible for the results and documentation of the work, and reviewers are expected to apply the same level of diligence as when no such tool was involved. Preparers, for their part, said that the "black box" nature of some tools and the lack of consistent output raise questions about the auditability of certain AI-created output, and that some of them are testing individual use cases to understand the extent of controls that would be needed before the technology is more widely integrated in financial reporting.
Firms investing in these tools and preparers alike stressed that human supervision and review of the output remain important, and none of the control frameworks exempts a control because software operates it. An organization putting agents into approval flows should therefore expect its auditors to ask how the control's operation is evidenced, who reviews the exceptions, and how the records would survive the departure of the engineers who built the system. The workable posture is the one effective human oversight of agents already points to: verify routine approvals automatically against the systems of record, route mismatches and edge cases to qualified reviewers, and keep records a reviewer outside engineering can actually read.
How Pruvz records an approval
Pruvz is a business evidence layer for production AI agents, and an approval workflow is a natural fit for its evidence model. Pruvz does not grant or block approvals: the agent and the organization's own systems keep that authority. What Pruvz does is keep the record that makes the approval defensible. For each agent-granted approval it creates a supplemental, action-level evidence record that can be linked to the transaction in the approval system or ERP that owns it. That record captures the decision-time context the agent acted on, as approved evidence fields, references, or extracts subject to data-minimization and retention controls, preserves an immutable snapshot of the policy in force (the thresholds and rules the approval was supposed to honor), records the decision and the executed action as ordered events, and then verifies the outcome independently in the system of record where the result lives, such as the billing platform or the CRM, instead of trusting the agent's report. The verification status comes from that read-back: it remains pending during a defined verification window and is finalized as verified, outcome mismatch, or verification failed. Mismatches land in a human review queue, and the reviewer's decision is appended to the same append-only record without rewriting the original result. You can walk through a real evidence packet field by field, and see how independent verification works in the system-of-record verification guide.
Pruvz is non-blocking by design: it records and verifies off the agent's critical path, so the approval flow keeps its speed. It does not replace the approval system, the ERP, or the other systems of record, and it does not by itself establish compliance with Section 404 or any other control requirement; it supplies the decision-time evidence and the independent outcome verification that a defensible approval record depends on.
Today, the demonstrated read-back runs against live Stripe and HubSpot test environments. ERP and approval-flow connectors are planned; they will be developed with founding design partners around the approval workflows selected with them, and the design-partner program is open now. If your agents approve invoices, expenses, discounts, or credits, and the record of those approvals is an "approved" flag in a database, that is exactly the gap a design-partner conversation is built to map. For the wider picture of what the record should contain, start with what business evidence for AI agents is.
This article is general technical and governance information, not legal advice. Your obligations depend on your jurisdiction, industry, and specific use case.