No. 127 / 339
What changes for scientific research with AI?
The shift
Synthesizing the entire published literature, generating candidate hypotheses, writing analysis code, and drafting papers all go from scarce, trained-researcher hours to abundant, near-instant, near-free output. What doesn't move: running the physical experiment, and knowing whether a result is actually true rather than merely plausible and well-written.
The axioms
- Generating novel, testable hypotheses requires scarce expert cognition and creative insight built over years of training.
- Knowing the full existing literature is gated by scarce reading time, which is why rediscovery and siloed work happen constantly.
- Producing data — running the experiment, collecting the sample, observing the system — requires scarce lab time, equipment, reagents, field access, or subjects.
- Peer review is the field's verification layer: an accountable expert vouches for correctness before a claim enters the record, gated by scarce reviewer time.
- The PI-led lab hierarchy exists because expert direction-setting is scarce relative to the execution labor needed to generate data, so a few scarce minds direct many hands.
- Junior researchers build scientific judgment by personally doing the slow grunt work — running the gel, writing the script, debugging the pipeline — under a mentor's eye.
- The paper is the unit of scientific currency because writing findings up clearly, and having them checked, was expensive, so results get batched into discrete, citable units.
- Replication is rare and low-status because it's expensive and unrewarded relative to novel work, so published results are trusted by default until someone spends scarce time to check.
- Funding is rationed by scarce grant-committee attention, which is why proposal-writing and review exist as a costly filter on who gets to do the expensive part (running experiments).
- Credit and career advancement (authorship order, tenure cases) are earned by demonstrated individual contribution to scarce intellectual output.
Invalid axioms
- Knowing the full literature is gated by scarce reading time. A model can ingest and cross-reference the entire published corpus of a subfield in minutes, in any language, including papers a human would never have found. The habit-trap: labs still budget months for "getting up to speed" on a literature, and reviewers still treat "I wasn't aware of that prior work" as a forgivable gap rather than a solved problem.
- Generating a first hypothesis is the hard, scarce part of a research program. Models can now propose dozens of plausible, literature-grounded hypotheses on demand, and rank them against known mechanisms. The habit-trap: grant timelines and PI bandwidth are still structured as if hypothesis generation is the bottleneck, when it's now the cheapest step in the pipeline.
- Writing the paper is expensive, so findings get batched into discrete publishable units. Drafting, formatting, and even generating figures from results is now fast and close to free. The habit-trap: the field still prices career advancement in paper count, incentivizing volume in exactly the step that no longer costs anything to produce.
- Analysis code and statistical pipelines are hand-built by whoever ran the experiment. Models write, debug, and adapt analysis code faster than most researchers can, including translating between languages and stats packages. The habit-trap: labs still hire and train for "can code the pipeline" as a differentiator when it's rapidly becoming table stakes anyone can get from a prompt.
- Grant and manuscript reviewers must personally read and synthesize every submission. A meaningful share of review-assist and first-pass screening is already AI-mediated at major venues, whether disclosed or not. The habit-trap: review is still scheduled and credited as if the reviewer's scarce reading time is what's being spent, while the actual scarce resource — judgment on whether the claim holds — hasn't been re-priced.
Unchanged axioms
- Producing new data still requires scarce physical or transactional action. Running the wet-lab experiment, collecting the field sample, recruiting the human subjects, operating the telescope or particle detector — none of this is a token-generation problem. AI can plan and even help automate parts of it (lab robotics, protocol optimization), but the physical throughput ceiling hasn't moved at anywhere near the rate synthesis has.
- Someone accountable must stand behind the claim that a result is true. A model can generate a confident, well-cited, textbook-plausible claim that is wrong — hallucinated citations, misapplied statistics, non-existent effects — and it can't be named as the responsible party when that happens. Authorship, retraction, and institutional consequences still attach to a human or a lab.
- Judgment on what's worth doing next, under real uncertainty, stays human. Choosing which of a hundred AI-generated hypotheses is actually worth the scarce lab time to test — given cost, feasibility, career risk, and scientific taste — is not a pattern-match problem. It's a bet made by someone who owns the consequences of being wrong.
- Ground-truth verification of a novel, high-stakes result is still scarce and slow. Replication, independent data collection, and adversarial scrutiny of a surprising finding take real time and real resources regardless of how fast the write-up was produced. A model can flag inconsistencies and suggest checks, but it can't independently confirm a physical or biological fact — only another measurement can.
- Trust between collaborators, and a mentor's judgment about a trainee's readiness, doesn't digitize. Deciding whether a junior researcher is ready to run independently, or whether a collaborator's data can be trusted, rests on a track record built over time — not on the fluency of an AI-assisted output.
New axioms
- When hypotheses and drafts are abundant, who filters signal from a much larger pile of plausible-sounding but untested claims? If generating a candidate hypothesis or a full manuscript draft costs nothing, the volume entering the pipeline can exceed what any human reviewer, editor, or funding committee can meaningfully triage — the bottleneck moves downstream to filtering, and nothing has been resourced to do that at the new scale.
- How does a field detect fabricated or subtly wrong AI-assisted results at the volume they can now be produced? Confident, fluent, citation-studded text that's factually or statistically wrong is the default failure mode of the technology generating it — and peer review was built to catch human error and human fraud at human production rates, not this.
- If junior researchers no longer have to do the slow manual grunt work to get a result, how do they build the judgment that used to come from doing it themselves? The apprenticeship model assumed struggling through the pipeline by hand was where scientific intuition got built. If a model does that step for them, it's untested whether judgment develops some other way or just doesn't.
- What happens to the incentive to run expensive, slow, real experiments when a model can generate a plausible simulated result or a compelling literature-grounded prediction instead? As synthetic data, simulation, and AI-predicted outcomes get cheaper and more persuasive, the temptation to substitute them for the actually scarce thing — real physical measurement — grows, especially under funding and publication pressure.
- Who is accountable when a discovery, hypothesis, or analysis was substantially AI-generated, and how does credit (authorship, tenure, prizes) get assigned? The credit system assumes a traceable line from a scarce human insight to the result. When the insight step is AI-assisted or AI-originated, institutions don't yet have a rule for who gets the credit, and who owns the fault if it's wrong.
- Does the paper remain the right unit of output when it can be produced far faster than the community can read or check it? If publishable units keep flowing at pre-AI cadence assumptions while production cost collapses, the literature itself risks becoming a volume problem before it's a knowledge problem.
Where it breaks
"Writing the paper and generating hypotheses were the scarce, hard parts of research" (invalid) collides directly with "ground-truth verification is still scarce and slow" (still holds): the pipeline now produces manuscripts and hypotheses far faster than experiments can be run or results replicated to check them. Journals and grant committees still schedule review as if submission volume tracks the old, slow production rate — it doesn't, and the backlog of unverified claims grows exactly where nobody resourced the check.
A second collision: "credit is earned by demonstrated individual contribution to scarce intellectual output" (invalid, once hypothesis and drafting labor is abundant) meets "who's accountable when an AI-assisted result is wrong" (new). Tenure committees and journals still ask "whose idea was this," but can't yet answer "who's liable if it's false" — the field is rewarding authorship on the old scarcity while accountability for AI-assisted claims has no owner at all.
Related axioms
Research
Who's accountable when a product decision is made on synthetic-user data that turns out not to reflect real users?
Research
Is hypothesis generation still a scientist's job when AI systems can propose and rank novel hypotheses themselves?
Research
How does lab structure change when one PI plus AI agents can do the throughput that used to require five postdocs?
Research
How do we measure research team impact when "insights delivered" is no longer a scarce output?
Research
Should research ops still gatekeep access to real participants now that synthetic panels are the default first pass?
Research
Is the paper still the right unit of scientific output when AI can generate them faster than humans can read them?
Other axioms
Society
What changes for communication and collaboration with AI?
Engineering
What changes for site reliability engineering with AI?
Healthcare
Is the pharmacist's checking role obsolete when AI verifies interactions and dispensing, or does accountability keep them?
Engineering
What shifts in accountability when an autonomous agent, not a human, executes the remediation?
Retail
What changes for retail with AI?
Legal
Should a judge rely on AI risk-assessment and sentencing tools, and who owns the decision?