No. 272 / 339
Do AI early-warning scores replace the nurse's read of a deteriorating patient, or just add alarm fatigue?
The shift
Continuous risk scoring — reading vitals, labs, trends, and nursing notes to flag a patient who is sliding before it's obvious — goes from a scarce, intermittent human act (a nurse pausing to integrate the picture during rounds) to abundant, always-on, and near-free. What a good nurse used to do a few times a shift, a model now does every few minutes on every bed at once.
The axioms
- Catching deterioration early is gated by scarce nursing attention — a nurse can only integrate the full picture on so many patients, so often.
- The nurse's read is holistic: it fuses the numbers with things not in the chart — skin, breathing, alertness, "something's off."
- An alert is worth escalating because alerts are relatively rare and effortful to produce, so one firing means someone judged it worth firing.
- Escalation is an accountable human act — a named clinician decides to call the rapid-response team and owns that call.
- Someone with a license must own the decision to act or not act on a warning sign.
- The scarce failure mode is the miss you never saw — a subtle sign the overloaded nurse didn't have time to register.
Invalid axioms
- Catching early deterioration is gated by scarce nursing attention. A model watches every patient's trend continuously and flags the slide hours before a busy nurse would circle back to that bed. The habit-trap: escalation protocols and staffing ratios are still built around intermittent human surveillance — as if the constraint is "a nurse noticing," when the constraint has moved to "someone acting on the thousand things now noticed."
- An alert means a human judged it worth firing. Alerts used to be scarce because producing one cost attention; now they're abundant and cheap to generate. The habit-trap: clinicians still treat an alert as carrying an implicit human judgment behind it, and interfaces still fire each score as if it were a considered call rather than one probabilistic read among thousands per shift.
Unchanged axioms
- The nurse's holistic read and physical assessment stay scarce and human. Mottled skin, a change in the smell of a wound, work of breathing, the patient who is suddenly quiet or confused — much of this never reaches the model as an input, and the bedside integration of it is not a token-generation problem. A score reads what's charted; the nurse reads the patient.
- Escalation is an accountable act a licensed human must own. Deciding to call the rapid-response team, override a reassuring score, or hold off — and answering for that decision — does not transfer to a model. Liability and license stay with a named clinician regardless of how good the score gets.
- The human who catches the false-negative is the scarce safeguard. The dangerous patient is the one the score reads as fine because the deterioration doesn't match its training distribution — a rare presentation, an atypical trajectory, a signal the model wasn't built to weigh. Noticing that the reassuring number is wrong is exactly the novel-ambiguity judgment models are weakest at, and it's where harm is worst.
- Being confidently wrong is unusually dangerous here. A score that is right-on-average but catastrophically wrong on an outlier causes direct physical harm, not a redo. This raises the bar on verification rather than lowering it — the opposite of most domains where cheap plausible output is a net win.
New axioms
- When alerts are abundant, alarm fatigue desensitizes the nurse to the one that matters. A ward that fires hundreds of scores a shift trains its nurses to dismiss them; the signal that would have saved a life arrives inside the same stream everyone has learned to tune out. The scarce act moves from noticing deterioration to designing an alert load a human can stay responsive to.
- Automation bias erodes the read that's supposed to back the score up. A score that's usually right teaches nurses to trust it on the rare miss — the reassuring number quietly substitutes for the independent assessment that was meant to catch it being wrong. The safeguard in STILL HOLDS decays precisely because the tool is good, and no one has designed for keeping the human read sharp when the machine is reliable enough to lean on.
- Accountability blurs when the score was wrong or was ignored. If a nurse escalates on a bad alert and floods the response team, or doesn't escalate because the score said fine and the patient crashes, who answers — the nurse, the ordering physician, the vendor, the hospital that set the threshold? Malpractice frameworks assume the deciding mind and the accountable party are the same; a score sitting between them isn't yet resolved.
- A model optimized to be right-on-average will be systematically wrong on the outlier, and the outlier is the patient who dies. Aggregate accuracy — the number vendors and procurement optimize for — is the wrong target when the cost is concentrated on the rare miss. Nobody yet owns the question of whose deaths are acceptable inside a "highly accurate" score's error bar.
Where it breaks
"An alert means a human judged it worth firing" (invalid) collides with "alarm fatigue desensitizes the nurse to the one that matters" (new): wards deploy continuous scoring on the old assumption that an alert is a considered signal, so they fire every one — and in doing so destroy the responsiveness that made alerts useful. The tool works by generating abundance; the workflow around it still treats each alert as scarce, and the gap is filled by nurses learning to ignore it.
A second collision: "catching deterioration is gated by nursing attention" (invalid) meets "automation bias erodes the read that backs the score up" (new). Systems buy scoring to cover the gaps in human surveillance, then quietly rely on the human read as the backstop for the score's misses — while the same deployment trains that read to defer to the score. The safeguard and the thing it's meant to safeguard are being eroded by the same tool, and staffing models still assume the nurse's independent assessment is fully intact.
Related axioms
Healthcare
What changes for medicine and healthcare with AI?
Healthcare
What changes for clinical trials with AI?
Healthcare
When AI flags every caries on the radiograph, does the dentist's job shift from diagnosis to defending against over-treatment?
Healthcare
If AI generates a competent meal plan for free, is the registered dietitian's value the clinical-risk catch rather than the plan?
Healthcare
What changes for drug discovery and pharma R&D with AI?
Healthcare
What changes for elder care with AI?
Other axioms
Media
What changes for chefs and professional kitchens with AI?
Media
What changes for film with AI?
Product Design
When any client can generate a photoreal room in minutes, what's left of the interior designer — taste, or the buildable/spec'd/priced reality behind the render?
Industries
Does the human operator's judgment stay essential, or does AI autonomy make the "human in the loop" symbolic?
Engineering
Who owns correctness when review volume outpaces human reviewer bandwidth and code starts merging unread?
Engineering
If anyone can query data in plain English, is the data analyst's job the SQL or knowing which question is worth asking?