No. 126 / 339

Is replication/verification now the scarce, valuable skill that hypothesis generation used to be?

The shift

LLMs make hypothesis generation abundant — pattern-matching against the literature to produce plausible mechanisms, novel angles, and testable claims is now cheap, fast, and high-volume. What they don't make abundant is ground-truth verification: confirming a result actually replicates, an effect is real and not p-hacked or a pipeline bug, and a claim survives contact with messy real-world data. That gap is now the bottleneck.

The axioms

  • Generating a novel, plausible hypothesis is hard and rare — it rests on scarce domain expertise and creative synthesis across a fragmented literature.
  • Peer review and citation count are reasonable proxies for a claim being true — they rest on review capacity being scarce enough that gatekeeping filters signal from noise.
  • Running an experiment / replication is expensive relative to having the idea — it rests on lab time, compute, and human hours being scarce, so replication got skipped and idea generation got rewarded.
  • A researcher who can generate many hypotheses is more valuable than one who checks a few carefully — rests on hypothesis supply being the limiting reagent in the pipeline.
  • Novelty is what gets funded, published, and cited — rests on attention and publication slots being scarce, so novelty (not correctness) becomes the currency that survives selection.
  • A single well-cited paper or benchmark score is trustworthy evidence — rests on the cost of independently re-deriving or re-running it being high enough that nobody bothers.

Invalid axioms

  1. Generating a novel, plausible hypothesis is hard and rare. LLMs can produce dozens of mechanistically coherent, literature-grounded hypotheses per hour, across domains a single human couldn't cover. The scarcity (domain synthesis + creative recombination) is flipped — this is exactly the "pattern-match against everything ever written" capability. Habit-trap: fields still reward and hire for idea generation (grant scoring on "novelty," PhD training built around "find your own question") as if this were the bottleneck, when it's now the cheapest part of the pipeline.
  2. A researcher who can generate many hypotheses is more valuable than one who checks a few carefully. Once hypothesis supply is unbounded, the marginal hypothesis is worth close to nothing until verified. Habit-trap: compensation, prestige, and first-authorship norms still track "who thought of it" rather than "who confirmed it holds up," so incentives lag the actual scarcity.
  3. Novelty is the currency that survives selection. When anyone can generate ten novel-sounding claims before lunch, novelty stops being a scarce, honest signal of anything — it's now the artifact LLMs produce most easily and most confidently, including when wrong. Habit-trap: journals, conference program committees, and funding panels still score on novelty as a proxy for merit, rewarding exactly the abundant thing.

Unchanged axioms

  1. Peer review / citations as a truth proxy erodes but the underlying need for gatekeeping doesn't. LLMs can write reviewer-caliber critiques and generate plausible-sounding replication reports, but they can't actually run the wet-lab experiment, can't be held accountable when they rubber-stamp a fabricated result, and can't independently gather new physical or transactional evidence. Verification-as-truth is still bottlenecked on physical action and accountable judgment, not just text synthesis.
  2. Running a real experiment or replication study is still expensive relative to having the idea. Compute-only fields (some ML, simulation, theory) are seeing this compress fast — replication is now partly automatable via re-running code and re-deriving math. But anything requiring lab work, field data, human subjects, or physical instruments still runs on scarce time, materials, and access. This is the sharpest fork in the audit: how "STILL HOLDS" this is depends entirely on whether the field's evidence is text/code-native or physical-world-native.
  3. Someone has to be accountable when a claim is wrong. An LLM can flag "this effect size looks inflated" or "this p-value pattern resembles p-hacking," but it carries no liability, can't be barred from future funding, and doesn't have a reputation to lose. The verifier's value isn't just running the check — it's being the named, answerable party who staked their credibility on the result. That's a trust/accountability scarcity AI doesn't touch.

New axioms

  1. Who verifies the verifier when AI can generate a fake replication as convincingly as a real one? If an LLM can produce a plausible-looking replication write-up, statistical reanalysis, or even synthetic data that passes surface checks, "a replication exists" stops being sufficient evidence — the field needs a way to verify that verification itself wasn't AI-confabulated, and no consensus mechanism for that exists yet.
  2. Hypothesis volume now exceeds verification capacity by an order of magnitude the field hasn't priced in. If generation is nearly free and verification still requires scarce lab time or accountable human judgment, the backlog of untested claims grows monotonically — the question isn't "is verification valuable" but "how does a field triage which of 10,000 plausible AI-generated hypotheses deserve the scarce verification slot."
  3. Verification itself is becoming partially automatable (code re-execution, statistical re-derivation, data pipeline audits) — but unevenly across fields. In compute-native domains this could flip verification from scarce to abundant too, within a few model generations, collapsing the very premise of this audit for those fields specifically while leaving it fully intact for wet-lab/physical-world science. Anyone building policy or incentives around "verification is the new scarce skill" needs to specify which kind of verification, because the automatable and non-automatable kinds are diverging fast.

Where it breaks

Funding and hiring still score "novelty of hypothesis" (INVALID #1, #3) at the same moment AI-generated hypothesis volume is outstripping the field's actual verification capacity (NEW #2) — the more a program rewards novel claims, the faster it manufactures a backlog nobody can check, and the reward structure never adjusts because novelty is still what's visible and legible to a review panel.

A second collision: fields are starting to treat "someone ran a replication" as sufficient evidence again (STILL HOLDS #1's proxy function), right as AI makes it cheap to generate a replication report that looks rigorous but wasn't independently or physically verified (NEW #1). The old proxy (a replication was published) breaks exactly when the thing it's supposed to filter for — costly, hard-to-fake effort — stops being costly to fake.

Related axioms

Other axioms