In December 1938, McKesson & Robbins, one of America's most respected drug companies, turned out to be carrying roughly $20 million in assets that never existed. The inventory sat in Canadian warehouses that weren't there. The receivables came from customers who weren't real.
The auditors didn't break the rules. They examined the paperwork management handed them, because observing inventory and confirming receivables with customers weren't required at the time. The audit met the standards of its day. The standards were what failed.
The reforms that followed came down to two questions: Did anyone examine the actual thing? And who checked the people who produced the evidence? Nearly 90 years later, those same two questions sit at the center of AI assurance.
Where AI Assurance Stands Today
AI assurance is roughly where auditing was in 1938: practice is ahead of any agreed standard. Methods, certifications, and auditors exist, but no one has settled what an AI audit actually has to examine.
California has started with the auditor. On September 9, 2026, Governor Newsom signed SB 813 and AB 1405. SB 813 creates a voluntary state designation for independent verification organizations, with criteria due by January 1, 2028. AB 1405 creates an AI Auditor Registry by January 1, 2029, for audits that state law requires. Neither law requires an organization to be audited.
Independence is being defined. What an audit has to examine is still open, and that gap is what your own framework has to fill. You'll sit on both sides of it: building the evidence for your own AI decisions, and being handed someone else's evidence and asked to trust it.
Governance, Assurance, and Audit: Where the Responsibility Lives
These three terms get used interchangeably, but they do different jobs.
- Governance decides how much authority AI is given, who approves it, the limits it operates within, and how it's measured. It's set before the system acts and revisited as risk changes.
- Assurance is the linked set of activities across an AI system's lifecycle that produce the evidence an independent party would need to have justified confidence the system is doing what was claimed.
- Audit is an independent, formal, point-in-time examination of whether the system conformed to what governance approved, using the evidence assurance produced.
Put simply: governance decides what AI may do, assurance produces the evidence that it did, and audit independently tests that evidence. You can't truly audit a neural network directly, so you build a broader framework of assurance around the entire AI system.
Decision-making in an AI-centric organization: Input governance (compliance and data risk) and output governance (ethics and AI risk) feed a single Data, Analytics & AI Leadership Committee. Below it, the Data, AI, and Decision Stewardship Committees test results for accuracy, traceability, context, and ethics before they reach the business.
Step 1: Map the Assurance Chain
AI assurance is a chain, not a checklist. Each link consumes what the one before it produced: use case approval, the Decision Record, build and test, risk-reward sign-off, run-time logs and monitoring, and independent audit.
Three artifacts in that chain matter most. The approved use case answers whether the system should exist and how much authority AI gets. The Decision Record answers what the system was specified to do. The audit evidence answers whether it did that.
They sit on the critical path: no use case, nothing to specify; no record, no criteria to test; no evidence, nothing to attest. They're also where industry guidance stops. Certifications like ISO/IEC 42001 and SOC 2 cover the organization, not individual decisions.
Authority is set at approval, decision by decision, along a spectrum: a human in the loop approves each decision, on the loop monitors and can intervene, over the loop sets boundaries and reviews, or is out of the loop entirely. For each decision, answer four questions: who may formulate the outcome and act on it, who owns the process, what triggers a change, and what would revoke the authority.
Step 2: Write the Decision Record
The Decision Record is where requirements become proof. It has six fields: source(s) consulted, data fields used, model and version, reason code, scope flag, and human accountability. Written before build, it specifies what the system may draw on, decide, and say. Referenced after, it proves the system did exactly that.
Take a credit decline. The reason code field holds an approved list of decline reasons. Outputs must map to that list, exceptions go to a queue, and each reason is stored with its prompt, response, and any override. The wording has to hold up externally, too: "Debt-to-income ratio: 42%. Approval threshold: 36%" tells an applicant what to fix. "Unable to approve your request at this time" doesn't.
On paper alone, the record assures nothing. It has to be configured into the system.
Step 3: Test Against the Record, Independently
An AI audit tests three things: whether the configuration matches the record, whether run-time logs show it was followed, and whether humans reviewed and acted as designed. Then it tests the record itself: is it complete, approved at the right level, and still current? Sampling the logs against the record is the AI equivalent of going to the warehouse.
Independence matters as much as it did in 1939. The Decision Steward produces the evidence; internal audit tests it and reports to the board, not to the steward. Audit never approves what it later audits.
If you bought the system, its decisions are still yours to prove. Require model and version change notices, comparable run-time logs, retention for as long as a decision can be challenged, and independent test access. If you can't get the evidence, grant less authority.
Where to Start
Start with one use case and one flow, not all of them. Pick one AI use case, answer the qualification and delegation questions, write its Decision Record, then configure at least one field and agree on what an independent party would test.
Then ask yourself: if an outside party asked you tomorrow to prove your highest-stakes AI decision, could you show the use case was approved, what the system was specified to do, and that it did it? Would they look at the warehouse, or at the paperwork?
Want to see where your organization stands? Take FSFP's 3-minute AI Readiness Assessment for instant results, and watch the full AIGov session here to bring ideas back to your team.
