No. 337 / 339
What changes for clinical trials with AI?
The shift
Protocol design, patient matching and recruitment, site selection, ongoing data monitoring, and statistical analysis all move from scarce, specialist-hour work to abundant, fast, near-free output — and synthetic control arms and disease-progression simulation get good enough to stand in for some real-world data. What doesn't move: dosing real humans and observing what actually happens in their bodies, which is the one thing a trial exists to find out and the one thing a model can only predict, not confirm.
The axioms
- Designing a protocol — endpoints, inclusion/exclusion criteria, statistical plan, sample size — requires scarce experienced medical and biostatistical judgment.
- Finding and enrolling eligible patients is slow and expensive because it requires reading unstructured records and matching them against complex criteria by hand.
- Choosing which sites and investigators to run at depends on scarce operational knowledge of where the right patients and reliable teams are.
- Monitoring incoming trial data for safety signals and quality problems requires scarce expert attention, so it's sampled and lagged, not continuous.
- Analyzing the results — the statistical work that turns data into a claim — is done by scarce trained biostatisticians on a hand-built pipeline.
- A control arm requires enrolling and following real patients who don't get the experimental treatment, which is costly, slow, and ethically fraught.
- Establishing whether a drug is safe and effective requires giving it to real humans and observing real outcomes over real time — this cannot be shortcut.
- A licensed, accountable sponsor and set of investigators must stand behind the trial's conduct and its claims to a regulator.
- Every enrolled patient must give informed consent and be protected from avoidable harm, overseen by an accountable ethics board (IRB/ethics committee).
- Regulators decide what evidence is acceptable, and that bar moves slowly and deliberately because the cost of being wrong is measured in patient harm.
Invalid axioms
- Designing a protocol is gated by scarce specialist judgment and months of drafting. Models can now generate a full draft protocol — endpoints, eligibility logic, statistical analysis plan, sample-size rationale — grounded in prior trials and guidance documents, in a fraction of the time. The habit-trap: sponsors still budget protocol development as a multi-month specialist bottleneck and price CRO design work as if the drafting itself were the scarce good, when the scarce part has moved to deciding whether the design is right.
- Finding eligible patients is slow because matching records to criteria is manual. Models read unstructured clinical notes at scale and match against complex inclusion/exclusion logic far faster than coordinators screening charts by hand — pre-screening cohorts that would have taken months. The habit-trap: recruitment timelines and site budgets are still built around manual chart review as the rate-limiter, even as that step collapses.
- Site selection depends on scarce, hard-won operational knowledge. Models can rank sites and investigators against enrollment history, patient-population data, and past performance far more cheaply than relationship-based intuition alone. The habit-trap: feasibility is still run as a slow, contact-driven exercise priced as specialist operational work.
- Data monitoring must be sampled and lagged because expert attention is scarce. Continuous automated monitoring of incoming trial data for anomalies, protocol deviations, and quality problems is now cheap enough to run on the full dataset in near-real-time rather than on periodic sampled review. The habit-trap: monitoring is still staffed and scheduled as periodic source-data verification, spending scarce human hours on a coverage problem that's now largely a compute problem.
- Turning trial data into an analysis requires scarce biostatistician hours on a hand-built pipeline. Models write, adapt, and document analysis code and generate the statistical outputs faster than most teams can hand-build them. The habit-trap: analysis is still resourced as scarce programming labor rather than as verification of a machine-produced result — which is where the scarce judgment actually now sits.
Unchanged axioms
- Establishing safety and efficacy still requires dosing real humans and waiting for real biology. Whether a drug works and whom it harms is a fact about bodies over time, not a token-generation problem. Simulation and synthetic data can sharpen a hypothesis and reduce how many patients you need, but they cannot produce the confirmatory fact — a model that predicts a drug is safe is making a plausible claim, and the entire purpose of a trial is that plausible isn't good enough. This is the load-bearing scarcity, and no capability trend on the horizon removes it.
- An accountable sponsor and investigators must stand behind the trial. When a trial harms someone or its data is wrong, a named human, institution, and legal entity are answerable to regulators, ethics boards, and courts. A model can draft, monitor, and analyze; it cannot be liable, hold a license, or be sanctioned. Accountability didn't get cheaper.
- Informed consent and patient protection stay human and irreducible. A real person is being exposed to real risk, and someone accountable — the investigator, the ethics board — must ensure they understood it and were protected. This is a duty owed between humans; it doesn't compress because the paperwork got faster.
- Ground-truth verification of a novel safety or efficacy result is still scarce and slow. A surprising signal — an unexpected adverse event, an efficacy result that's too good — has to be checked against the actual data and often the actual patients, adversarially and by people who own being wrong. A model can flag inconsistencies; it can't independently confirm that a biological effect is real. Confidently-wrong output is the default failure mode of the tools now doing the design and analysis, which raises the value of this check rather than lowering it.
- The regulatory evidence bar moves deliberately, and trust in it is what makes trials mean anything. The acceptability of any given evidence — including AI-designed protocols and synthetic control arms — is set by regulators weighing patient-harm risk, and that bar advancing is a slow, accountable process by design. (Fast-moving: regulatory acceptance of AI-assisted design, decentralized/AI-monitored trials, and synthetic/external control arms is advancing faster than usual as of mid-2026 — this is the axiom most likely to partially shift within a couple of years, and the specific acceptable use-cases are the thing to watch.)
New axioms
- Design and recruitment now outrun the irreducible clock of dosing real patients. When you can design a protocol and assemble an eligible cohort in a fraction of the old time, the trial's duration is set almost entirely by the biology — how long you must dose and follow patients to observe the outcome. Speeding up everything around the dosing window makes that window the dominant constraint, and the cost/timeline models that assumed design and recruitment were the long poles are now mis-forecasting where the time actually goes.
- Over-trusting synthetic and simulated evidence, especially under commercial pressure. As synthetic control arms and simulated outcomes get cheaper and more persuasive, the pull to substitute them for real enrolled patients grows — precisely where a real control arm is slow, expensive, and ethically awkward. The open problem is where simulated evidence genuinely reduces the human exposure needed versus where it quietly launders a plausible prediction into a claim that only real patients could have earned.
- Who is accountable for an AI-influenced safety miss. When an AI monitoring system fails to escalate a signal, or an AI-drafted protocol has an eligibility flaw that exposes patients to harm, the accountability chain has to land on a named human or entity — but the failure is now distributed across a tool, its vendor, and the people who trusted it. The rule for who owns that miss doesn't exist yet, and "the model got it wrong" cannot be where it stops.
- Verifying AI-produced trial analysis at the speed and volume it's now generated. If the statistical analysis and its narrative come out of a model, someone still has to confirm the numbers and the claim are right — and review was built to check human statisticians working at human speed, not to catch fluent, citation-studded, subtly-wrong machine output at the volume it can now be produced. Verification, not production, is the constraint, and it isn't resourced for the new rate.
- The ethics of AI-selected cohorts. When a model picks who is eligible and which sites to run at, it can inherit and amplify biases in the historical data — under-enrolling populations already underrepresented in trials, or optimizing for enrollment speed over generalizability. The open problem is auditing cohort selection for fairness and representativeness at the point the AI proposes it, before it silently shapes who a drug is ever tested on.
Where it breaks
"Design and recruitment are the long, scarce poles of a trial" (invalid) collides with "establishing safety and efficacy still requires dosing real humans and waiting for real biology" (still holds): teams now compress everything up to first-dose, then discover the calendar is governed by the follow-up window they can't shrink. The temptation that grows in exactly that gap is NEW problem 2 — substituting synthetic or simulated evidence for the slow real thing — because it's the only remaining lever that appears to move the timeline, and the pressure to pull it is strongest precisely where the evidence is weakest.
A second collision: "monitoring and analysis are scarce human labor" (invalid) meets "who's accountable for an AI-influenced safety miss" (new). Sponsors are moving monitoring and analysis onto automated systems to capture the speed and coverage, while the accountability framework still assumes a human was in the loop reading the data — so responsibility for a missed signal is being quietly transferred to tools that cannot hold it, and no one has re-assigned the ownership that moved.
Related axioms
Healthcare
What changes for medicine and healthcare with AI?
Healthcare
When AI flags every caries on the radiograph, does the dentist's job shift from diagnosis to defending against over-treatment?
Healthcare
If AI generates a competent meal plan for free, is the registered dietitian's value the clinical-risk catch rather than the plan?
Healthcare
What changes for drug discovery and pharma R&D with AI?
Healthcare
What changes for elder care with AI?
Healthcare
If ML triage out-performs the call-taker's protocol, does accountability for the dispatch decision move to the model — and who's liable for the under-triage that kills?
Other axioms
Retail
Does online merchandising and category management still need a human curator when AI personalizes the storefront per shopper?
Media
If AI generates the docs from the code, is the technical writer the writing or the ownership of whether the docs are true?
Industries
What changes for agriculture with AI?
Education
What changes for vocational education with AI?
Marketing
Is SEO dead as a discipline now that AI answers (search AI overviews, chatbots) replace the results page it was built to rank on?
Engineering
Who owns a CI/CD pipeline once AI agents can write, debug, and modify it directly instead of a dedicated DevOps engineer?