No. 121 / 339

Is the paper still the right unit of scientific output when AI can generate them faster than humans can read them?

The shift

Writing up a study — literature synthesis, methods prose, results narrative, framing — goes from scarce (months of trained researcher time) to abundant (minutes, near-zero cost, unbounded volume). What doesn't move with it: whether the underlying claim is true, reproducible, or worth anyone's attention. AI makes the write-up cheap; it doesn't make the finding correct.

The axioms

  • The paper is the unit of record because writing one is expensive. Expensive drafting was a natural filter — only work someone thought was finished and defensible got typed up. Rests on drafting being scarce.
  • Peer review works because reviewer time is the bottleneck, so only a manageable number of papers need reading. Rests on submission volume being naturally capped by how slow writing is.
  • A paper's length and structure (abstract, methods, results, discussion) signal how much labor went into the work. Rests on the correlation between effort-to-write and effort-to-produce-the-result.
  • Citation count and publication count are usable proxies for contribution. Rests on both sides — writing and reading — being similarly slow, so volume tracked substance.
  • A human author is answerable for what the paper claims. Rests on accountability being tied to a person who did the work and put their name on it.
  • Novelty and correctness get judged by a small number of expert readers before the result enters the record. Rests on expert judgment being scarce and therefore concentrated in a gatekeeping step.
  • Reproducing a result requires redoing scarce, slow human labor (rerunning experiments, re-deriving proofs, re-examining data). Rests on verification being roughly as expensive as the original work.

Invalid axioms

  1. The paper is the unit of record because writing one is expensive. AI makes drafting free and instant — full papers, complete with plausible citations and figures, generated in minutes. The scarcity that made "finished enough to write up" a meaningful filter is gone. Habit-trap: journals, tenure committees, and funders still size review pipelines and publication counts as if submission volume were self-limiting by writing cost. It isn't anymore — submission counts are already climbing as a direct result.
  2. A paper's length and structure signal how much labor went into the work. Length and formal completeness used to correlate with effort. AI can produce a fully-structured, fluent paper around a thin or fabricated result as easily as around a rigorous one. Habit-trap: reviewers and readers still use polish and structure as a quality heuristic — exactly the signal AI is best at faking.
  3. Citation count and publication count are usable proxies for contribution. When both writing and (soon) much of the reading/summarizing is AI-mediated, volume stops tracking substance at all — it can be gamed by generating more, denser, more citable-looking output. Habit-trap: hiring, promotion, and funding committees still lean on paper counts and citation metrics as the load-bearing evidence of a researcher's output.

Unchanged axioms

  1. A human (or a named, accountable entity) must be answerable for what the paper claims. A model can generate a claim; it can't be held liable, retracted with consequence, or lose standing in a field. Someone's name, reputation, and career still have to be on the line for the claim to mean anything in the scientific record.
  2. Novelty and correctness get judged by expert readers before a result enters the trusted record. AI can flag inconsistencies, check statistics, and summarize prior work fast, but judging whether a result is actually new, actually matters, and actually holds up under adversarial scrutiny in a genuinely novel domain is still a human judgment call — current models are fluent, plausible, and not reliably correct on frontier claims. This is the axis most likely to erode as verification tools improve; recheck it as AI-assisted review tooling matures.
  3. Reproducing a result requires redoing the underlying work, not just reading about it. AI can write and rewrite a paper about an experiment infinitely; it doesn't run the wet lab, weigh the sample, or observe the physical world (except where the science is purely computational). For any field with a physical or transactional component, ground-truth verification still requires action AI can't take.

New axioms

  1. When papers can be generated faster than any human — or team of humans — can read them, what unit does the record use instead? If output volume is unbounded, the paper-as-atomic-unit stops being a viable filter for attention. The field has no agreed replacement: not "read every paper," not "trust citation count," not "let AI summarize everything and hope the summarizer is right."
  2. Who verifies at the new volume, and what do they verify against? Confidently-wrong papers are now cheap to produce at scale. If review capacity stays human and roughly fixed while submissions become unbounded, either review quality drops, or a much larger share of the record goes effectively unreviewed — and nobody has decided which.
  3. Does provenance — knowing how much of a paper's reasoning is AI-generated versus human-verified — become a first-class part of the scientific record? Right now there's no standard for disclosing or auditing this, but the incentive to blur it (more papers, faster) is already live.

Where it breaks

Volume filtering by drafting cost is gone (INVALID #1), but the field still has no substitute filter for deciding what's worth a scarce expert's attention (NEW #1) — journals are seeing submission counts rise while reviewer pools stay flat, and nobody has redesigned the funnel, they've just made reviewers do more with AI-assisted triage tools of uncertain reliability.

Citation/publication count as a proxy for contribution is dead as a trustworthy signal (INVALID #3), right as the need for a real verification layer becomes urgent (NEW #2) — and the two are currently solved by the same actors under the same incentive: researchers and journals still get rewarded for more output, at the exact moment more output is the thing making verification harder.

Related axioms

Other axioms