No. 66 / 339
Does an audit still mean anything when AI drafted the work papers it's supposed to check?
The shift
Producing the work papers — reconciliations, sample testing, variance analysis, control narratives, first-pass conclusions — goes from scarce associate-hours to abundant, near-instant drafting. The audit's evidentiary trail can now be generated by the same class of system, at the same time, as the transactions it's meant to test.
The axioms
- The preparer and the reviewer are structurally separate, which is what makes review meaningful (scarcity of who has access to produce vs. check).
- Work papers are produced slower than they're reviewed, so review time is the natural check on volume (drafting is expensive; checking is comparatively cheap).
- A sampled subset stands in for the whole population because testing everything is too costly (exhaustive testing is scarce).
- The auditor's signature carries weight because it represents scarce, licensed, personally liable judgment (accountability is scarce).
- Professional skepticism is a human disposition applied by someone with nothing to gain from the client's numbers being right (independent judgment under incentive is scarce).
- An audit trail is trustworthy because each step was authored by someone who can be deposed, sanctioned, or fired (a traceable, accountable author is scarce).
Invalid axioms
- Drafting the work papers is the bulk of audit cost and time, so review naturally lags and samples. AI collapses drafting to near-zero cost and time — reconciliations, tie-outs, variance narratives, and first-pass testing can be generated for the full population, not a sample, in the time it used to take to draft one workbook. Habit-trap: firms still budget, staff, and bill audits as if drafting hours were the scarce resource, front-loading junior headcount on work that's now largely automatable, while review capacity hasn't grown to match the volume AI can now produce.
- Sampling is necessary because testing 100% of transactions is too expensive. Full-population testing was scarce because it required scarce human hours; AI makes exhaustive testing cheap. Habit-trap: methodologies still default to statistical sampling sized for a world of expensive testing, leaving firms under-using a capability that's now essentially free and over-relying on inference from a subset when the whole population is checkable.
Unchanged axioms
- The auditor's signature carries scarce, personal accountability. A model can draft a memo concluding revenue recognition is appropriate; it cannot be sanctioned by a regulator, sued by investors, or stripped of a license. Someone must still be the answerable party, and that person must still understand and stand behind the conclusion — not just the output that led to it.
- Professional skepticism under incentive pressure is a human judgment call, not a pattern match. AI is fluent at generating a plausible-sounding rationale for why a number is fine — that's exactly the failure mode that matters here, because a model has no stake in whether it's colluding with management's preferred answer. Detecting collusion, override, or a client steering the narrative requires a party structurally independent of the incentive, which AI is not.
- Novel, high-stakes judgment calls — going-concern, fraud risk, materiality on an ambiguous transaction — have no clean pattern to match against. These are exactly the situations where confidently-wrong is most dangerous and where there's no large corpus of "how this always goes" for a model to pattern-match to. This stays a human call.
- The audit's value to third parties (investors, regulators, lenders) rests on trust in an accountable institution, not on the artifact's polish. A well-formatted work paper was never the point — the point was that a licensed, liable party vouched for it. That standing doesn't transfer to a tool.
New axioms
- If AI drafts the work papers, what is the reviewer actually reviewing? Traditional review assumes the preparer did the analytical work and the reviewer checks judgment and completeness. When the preparer is a model, the reviewer may be checking fluent, well-formatted output that was never actually reasoned through — the appearance of rigor decoupled from the substance of it. Firms need a defined answer for what "review" means when the first draft has no human judgment embedded in it to review.
- Volume of testing no longer bottlenecks the audit, but verifying AI's own conclusions at that volume does. If a model can test 100% of transactions, someone still has to verify the model's flags and non-flags are correct — and doing that at full-population scale is a new capacity problem nobody staffed for. Cheap testing doesn't mean cheap verification of the tester.
- Independence and skepticism assumed a preparer with human incentives to detect; an AI preparer has no incentive at all, in either direction. That sounds safe, but it creates a new blind spot: a model trained on plausible patterns will produce a plausible-looking work paper even when the underlying transaction is genuinely anomalous, and nobody has defined who's responsible for catching "confidently clean" AI output that's actually wrong.
- Regulators and standard-setters haven't specified what "prepared by AI, reviewed by a human" needs to look like for it to satisfy audit standards. This is the fastest-moving piece — PCAOB, IAASB, and AICPA guidance on AI-assisted audit evidence is still being written, and any answer given today should be treated as provisional.
Where it breaks
Firms are cutting associate headcount and staffing structure on the assumption that AI-drafted work papers absorb the "expensive drafting" problem (INVALID #1) — but the new bottleneck is verifying AI's conclusions at the full population scale AI just made possible (NEW #2). Cutting the junior tier that used to build pattern-recognition skill through manual drafting removes the exact training ground that produces reviewers capable of catching a model's confidently-wrong output later. The field is solving the cost problem and quietly deleting the pipeline that produces the people who can still do the one thing that stays scarce: skeptical, accountable judgment.
Related axioms
Finance
What changes for finance and banking with AI?
Finance
What changes for insurance with AI?
Finance
Is the branch banking model dead when AI can handle account service, loan applications, and advice remotely?
Finance
What's a financial advisor for once portfolio construction and tax-loss harvesting are commoditized?
Finance
Is manual bookkeeping just dead now that reconciliation is free?
Finance
Should finance still "close the books" monthly if AI can close them continuously?
Other axioms
Healthcare
Who's accountable when a semi-autonomous surgical robot, guided by AI, is involved in a bad outcome?
Media
If AI generates the docs from the code, is the technical writer the writing or the ownership of whether the docs are true?
Media
If a client can train a LoRA on my back catalogue and generate "more in my style," what is an illustrator selling — images or authorship?
Legal
Is the paralegal role dead, or does it just move upstream into AI-output verification?
Marketing
Is guided onboarding still a human-led function when AI can walk new customers through setup conversationally?
Industries
What changes for mining and extractive industries with AI?