No. 227 / 339
Does "the adjuster reviews the outcome" still count as human oversight, or is it theater?
The shift
Producing a claims decision — reading the file, applying policy terms, drafting the rationale, computing the payout — goes from scarce adjuster hours to abundant and near-instant. Review capacity does not move with it. When the same human no longer produces the decision, "a human reviews the outcome" quietly detaches from any fixed amount of reviewer time per decision, so oversight can be nominally present while the seconds available to actually re-examine each decision collapse toward zero.
The axioms
- A human reviewing a claims decision before it goes out is meaningful oversight — because historically the reviewer was also the producer, and producing took real time, so review time scaled roughly with decision volume.
- Oversight is a control: it catches wrong or unfair decisions before they reach the claimant — because the reviewer has the time, the underlying information, and the standing to overturn the decision.
- Regulatory good-faith and fair-claims-handling duties are satisfied when a qualified human is in the loop — because "human in the loop" was a reliable proxy for genuine, individualized consideration.
- The insured retains recourse to a real, accountable decision-maker — because someone with authority actually engaged with the specific facts of their claim.
- A human signing off transfers accountability to that human — because signing meant having formed and owned a judgment.
Invalid axioms
- A human reviewing the decision before it goes out is, by itself, meaningful oversight. This rested on review time scaling with decision volume, which held only because the reviewer was also the slow producer. Once production is abundant and instant, a reviewer can be handed more decisions per shift than there are seconds to open them. The habit-trap: carriers count "% of automated decisions with human review" as an oversight metric and report it to boards and regulators as a control, while never checking it against the arithmetic — decisions per reviewer-hour — that determines whether review can be anything but a click.
- A human sign-off transfers genuine accountability to that human. Sign-off was a proxy for having formed a judgment, because forming one was the expensive step you couldn't skip. When the judgment is pre-formed by the model and the human's act is confirmation, the signature still transfers nominal liability but no longer certifies that a mind engaged with the facts. The habit-trap: process designers treat the sign-off field as the accountability mechanism, so the org believes it has a decision-owner when what it has is a name attached to an output nobody re-derived.
Unchanged axioms
- Genuine, accountable review on contested and complex claims requires scarce human judgment. A disputed liability call, an ambiguous exclusion applied to unusual facts, a claimant-credibility question — these have no ground truth to pattern-match against, and someone must decide and be answerable. Nothing here got cheaper; the constraint is real reviewer attention, which is exactly what volume-blind automation spends first. The audit's premise stands: on these claims, review is oversight or nothing. On routine, uncontested claims that match a clean pattern, "review" was arguably always light — the theater question bites hardest on the contested files that get swept into the same fast queue.
- The regulatory good-faith duty is a duty of genuine consideration, not of ceremony. Fair-claims-handling and good-faith obligations require that the specific claim was actually evaluated. A review step that certifies nothing does not discharge that duty just because a human touched the record — it exposes the carrier to bad-faith liability precisely because the control was represented as real. The duty rests on individualized judgment, which stays scarce; the paperwork of oversight does not satisfy it. (Whether regulators can detect hollow review is a fast-moving enforcement question — see Where it breaks.)
- The insured's recourse depends on a real decision-maker having authority and time to reconsider. Recourse is worth something only if the person on the other end can actually reopen and overturn the decision. That capacity — authority plus available attention — is scarce and did not scale with automated decision volume. A human who can only rubber-stamp cannot provide recourse, regardless of title.
- Accountability cannot sit with the model. A model can produce the decision and the rationale; it cannot be fined, sued, or hauled before a regulator. So the accountable party must remain human and institutional. The open problem is that this is now true by legal construction while being false in operation — the accountable human may not have meaningfully participated. The requirement holds; its satisfaction is what's in doubt.
New axioms
- The arithmetic gap between decision volume and reviewer-seconds has no owner. When one model produces in a shift what a department used to produce in a week, and the review headcount is flat or cut, the seconds-per-decision available to reviewers falls below what any genuine re-examination requires — often below the time to read the file at all. Nobody currently owns the metric "reviewer-seconds available per decision vs. reviewer-seconds a real review needs," so the gap widens invisibly behind a green "human-reviewed" indicator.
- Oversight that certifies nothing is being presented as the control. The dangerous case isn't the absence of review; it's review that is reported — internally, to regulators, to claimants — as the safeguard while being incapable of catching anything. This is worse than no review, because it manufactures false assurance and suppresses the demand for a control that would actually work. We must solve for distinguishing, on paper and in audit, review that can overturn a decision from review that can only confirm it.
- We haven't defined what non-rubber-stamp oversight actually requires. Real review of an AI-produced decision is a different job than producing the decision was: it needs the reviewer to see why the model decided, the inputs it relied on, a defensible reason to trust or distrust those inputs, enough time to re-derive the contested part, and authority and incentive to reverse it against the model's confident output. None of that is supplied by a sign-off field. Designing review that samples intelligently, escalates the contested, and measures its own catch rate — rather than reviewing everything nominally and nothing genuinely — is unsolved in most claims operations.
- Automation bias makes the human reviewer a weaker check than an unaided one. A confident, fluent, pre-formed model decision anchors the reviewer toward agreement; the human is now less likely to dissent than if they'd started from a blank file. So adding the review step can reduce scrutiny relative to no automation, while being counted as added oversight. We must solve for review designs that resist the pull of the model's confidence rather than inheriting it.
Where it breaks
"We report human-reviewed rates as our oversight control" (INVALID #1) collides with "the arithmetic gap between decision volume and reviewer-seconds has no owner" (NEW #1): the carriers automating decision production fastest are generating the largest gap between decisions issued and review capacity, and the same speed that creates the gap also inflates the review-rate number they present as proof it's under control. The metric goes up as the oversight it claims to measure goes to zero.
Separately, "a human sign-off transfers accountability" (INVALID #2) collides with "the good-faith duty is a duty of genuine consideration" (STILL HOLDS #2): carriers are relying on the sign-off to discharge a legal duty that the sign-off no longer evidences, so the fastest-automating claims shops are accumulating undetected bad-faith exposure on exactly the contested denials where the review was supposed to be doing the work — and whether that exposure ever surfaces hinges on regulators developing the ability to distinguish real review from theater, which today they largely cannot.
Related axioms
Finance
What changes for finance and banking with AI?
Finance
What changes for insurance with AI?
Finance
Is the branch banking model dead when AI can handle account service, loan applications, and advice remotely?
Finance
What's a financial advisor for once portfolio construction and tax-loss harvesting are commoditized?
Finance
Does an audit still mean anything when AI drafted the work papers it's supposed to check?
Finance
Is manual bookkeeping just dead now that reconciliation is free?
Other axioms
Engineering
Is manual test-case writing dead now that AI can generate test cases from a user story?
Education
What changes for higher education with AI?
Cybersecurity
How does a CISO's risk calculus change when both attackers and defenders run autonomous AI agents?
Society
Who captures the productivity gains — labor or capital — when AI makes execution abundant across the economy?
Government
What changes for government with AI?
Management
Do we still need the analyst who builds slides if the AI builds better slides faster?