No. 199 / 339

If ML triage out-performs the call-taker's protocol, does accountability for the dispatch decision move to the model — and who's liable for the under-triage that kills?

The shift

Acuity judgment from a phone call — deciding how sick or hurt a caller is, and what to send — goes from scarce (a trained dispatcher working a memorized protocol) to abundant, when a model that pattern-matches the call against far more prior cases can assign acuity in seconds. Research already shows ML triage cutting over-triage by roughly 15% against human call-takers, meaning it sends fewer unnecessary lights-and-sirens responses for the same detection of the genuinely critical. But the paramedic's physical response and the legal answerability for a fatal miss stay exactly where they were: human.

The axioms

  • Assigning acuity from a phone call is gated by scarce, trained human cognition working a protocol from memory.
  • Triage capacity scales with the number of trained call-takers you can staff — it's a human-headcount bottleneck.
  • The dispatch decision and accountability for it live in the same place: the person who made the call.
  • Getting help to the patient requires physical action — a crew, a vehicle, hands on the body.
  • When triage is wrong and someone is harmed, a named, licensed party has to answer for it.
  • Ambiguous, novel, or deteriorating calls need human judgment because there's no clean protocol branch to match.
  • The caller has to trust the voice on the line enough to disclose, comply, and stay calm.
  • Over-triage (sending too much) is the safe error; systems tune toward it because a false alarm is survivable and a missed critical is not.

Invalid axioms

  1. Assigning acuity from a phone call is gated by scarce, trained human cognition. A model can score the call against a far larger case base than any dispatcher holds in memory, faster and more consistently, and — on the over-triage metric — more accurately on average. The protocol memory that was the call-taker's core skill is now the commodity part of the job. The habit-trap: centers still staff, train, and certify around protocol recall as the scarce competency, and still measure a good call-taker by fidelity to the script rather than by the harder thing the model can't do — catching the call that doesn't fit the script.
  2. Triage capacity is a human-headcount bottleneck. If the model can assign acuity, the constraint on how many calls get a fast, consistent first read is no longer how many trained humans are on shift. The habit-trap: capacity planning, queue design, and "hold times acceptable because triage is slow" all still assume the scarce input is dispatcher-minutes, so centers keep rationing the wrong resource and under-invest in the resource that's now actually scarce — the human who reviews the ambiguous calls and owns the misses.

Unchanged axioms

  1. Getting help to the patient requires physical action. A crew, a vehicle, a defibrillator, hands doing compressions — none of it is a token-generation problem. Better triage changes who gets sent and how fast; it does nothing about the scarce resource of trained responders and units on the road. If anything, freeing capacity on the phone throws the constraint harder onto the physical fleet.
  2. A named, licensed party has to answer when triage is wrong and someone dies. Liability does not transfer to a model. An algorithm can't be sued, licensed, sanctioned, or made to testify. When under-triage kills, the accountable party is still the human — the dispatcher who accepted the recommendation, the medical director who approved the tool, or the agency that deployed it. This is a legal and institutional fact, not a capability gap, so it does not move as the model improves.
  3. Ambiguous, novel, and deteriorating calls still need human judgment. The atypical MI presenting as indigestion, the caller who minimizes, the situation changing mid-call, the case that sits between two protocol branches — these are exactly where pattern-matching against prior cases is weakest, because the signal that mattered wasn't in the pattern. A model tuned to be right on average is most exposed precisely on the calls that don't average out.
  4. The caller's trust is load-bearing. A caller has to disclose honestly, follow pre-arrival instructions, and stay calm enough to be useful. That runs on trusting the voice on the line — a competent, present human who can be held to account — not on the accuracy of a score the caller never sees.
  5. Over-triage is the deliberately safe error. The system tunes toward sending too much because a false alarm is recoverable and a missed critical is not. This asymmetry is a values choice, not an efficiency bug. A model optimized to reduce over-triage is being pushed against the direction the safety asymmetry points, which raises the bar on how a miss gets caught rather than lowering it. (Fast-moving: as models get more reliable, the defensible amount of over-triage to trade away will keep shifting — this calibration is not stable and shouldn't be frozen.)

New axioms

  1. Who is liable when the model under-triages and someone dies? A fatal under-triage now has at least four candidate owners — the dispatcher who deferred to the recommendation, the medical director who signed off on the tool, the vendor who built and tuned it, and the agency that chose the operating point. Malpractice and negligence frameworks were built for a world where the deciding mind and the accountable party were the same person. Shared human-model decisions break that before the law has caught up, and "the model recommended it" is not yet a settled defense or a settled fault.
  2. Automation bias when the model is usually right. A tool that beats humans on average trains dispatchers to defer to it. The better it performs, the more that deference is rewarded — right up until the call where it's wrong, which is now also the call where the human most-conditioned-to-defer is least likely to override. Better average performance can degrade the human backstop it depends on. Nobody has designed the workflow that keeps the human genuinely in the loop rather than nominally rubber-stamping.
  3. The rare catastrophic miss buried inside better average performance. A 15% cut in over-triage is a population-level number; it says nothing about the shape of the tail. A model can be better on average and still fail on a specific, rare, lethal presentation in a way no human would — or fail systematically on a subgroup underrepresented in its training cases. Averages hide correlated tail risk, and EMS has no established way to audit for the miss that's rare enough to vanish in the aggregate but lethal every time it lands.
  4. What trains the next generation of judgment if the model does the routine triage? Dispatchers built the instinct for the atypical call by working thousands of ordinary ones. If the model absorbs the routine volume, the pathway that produced the human who's supposed to catch what the model misses is untested at scale.

Where it breaks

"Triage capacity is a human-headcount bottleneck" (invalid) collides with "who is liable when the model under-triages and someone dies" (new). The efficiency case for the model is to let it carry more of the volume with fewer humans in the loop — but the only thing that can absorb liability is a human in the loop, and thinning that layer to capture the savings quietly moves the accountability for a fatal miss onto whoever is left, usually the individual dispatcher, without anyone deciding that on purpose.

A second collision: "assigning acuity is gated by scarce human cognition" (invalid) meets "automation bias when the model is usually right" (new). Centers deploy the tool because it out-performs the human — and that same out-performance is what erodes the human's willingness to override it on the rare call where overriding is the whole point of keeping a human there. The better the model gets, the more the human backstop is present on paper and absent in practice, and no one has resourced or measured the override as its own scarce, load-bearing act.

Related axioms

Other axioms