No. 117 / 339

Is hypothesis generation still a scientist's job when AI systems can propose and rank novel hypotheses themselves?

The shift

AI makes breadth of combinatorial search abundant — pattern-matching across a field's entire literature to propose large numbers of novel, mechanistically plausible hypotheses in minutes, and ranking them by fit to existing data. What stays scarce is knowing which of those hypotheses is worth the months of wet-lab or field time to actually test, and whether the ranking reflects real biological or physical plausibility versus literature-frequency bias.

The axioms

  • A scientist's expertise is measured partly by the breadth of prior work they can hold in mind while generating ideas — this rested on human recall and reading speed being scarce.
  • Novel hypotheses come from a trained expert noticing an unusual connection across disparate literatures — this rested on cross-domain synthesis being slow and rare.
  • Generating many candidate hypotheses is costly, so scientists generate few and commit early — this rested on ideation itself being expensive.
  • A hypothesis is only as good as the experiment that can falsify it, and someone has to design and run that experiment — this rests on physical action and lab access being scarce.
  • Ranking hypotheses by plausibility requires judgment about mechanism, prior probability, and what's actually testable given real-world constraints (funding, ethics, time) — this rests on contextual judgment, not just pattern-fit.
  • The person who proposes a hypothesis is accountable for the resources spent testing it and the claims made from the result — this rests on accountability being tied to a name.
  • Deciding which question is worth asking at all — which gap in a field matters — is a taste and priority call, not a search problem — this rests on scarce judgment about what's worth doing.

Invalid axioms

  1. Breadth of recall is a proxy for scientific insight. AI now reads and cross-references the full literature of a field (and adjacent fields) faster and more completely than any human can. The habit-trap: labs still treat the "well-read" postdoc or the person who "knows everyone's work" as the ideation bottleneck, and route hypothesis generation through them, when a model can already surface the same connections faster and generate more candidates than that person would in a week.
  2. Generating many candidate hypotheses is expensive, so you generate a handful and commit. AI collapses the cost of producing dozens or hundreds of plausible hypotheses to near zero. The habit-trap: grant proposals and thesis committees still expect a single early-committed hypothesis with sunk reasoning behind it, rewarding the appearance of foresight over what's now cheap — actually generating and comparing a wide candidate set before committing lab time.
  3. Cross-domain synthesis (spotting a mechanism from field A that explains a puzzle in field B) is a rare individual talent. This was scarce because no one person reads deeply across enough fields. Models with broad training now do this by default. The habit-trap: institutions still prize and specially fund "interdisciplinary" hires as the mechanism for cross-pollination, when the pollination itself is now a commodity — the scarce part has moved to knowing which cross-domain suggestion is real versus superficial pattern-matching.

Unchanged axioms

  1. Someone has to design and run the experiment that falsifies the hypothesis. AI can propose a mechanism; it cannot pipette, culture cells, run a particle detector, or recruit trial participants. Hypothesis generation and hypothesis testing are decoupled by this shift, but testing remains physical-world-scarce, so the bottleneck in most fields simply moves downstream rather than disappearing.
  2. Ranking hypotheses by real-world testability and cost is a judgment call, not a scoring function. A model's ranking is trained on textual plausibility and citation patterns — it doesn't know your lab's equipment, your funding cycle, or that a particular reagent has an 18-month backorder. Deciding which of 200 AI-generated hypotheses is worth six months of grad student time is a judgment about scarce resources the model doesn't have visibility into.
  3. Accountability for a claim stays with a named scientist. If a hypothesis leads to a retraction, a wasted grant, or a harmful clinical recommendation, a person answers for it — a co-author, a PI, a reviewer. AI proposing the hypothesis doesn't create a locus of accountability; it just means the human who selected and pursued it owns the consequences more fully, since "the model suggested it" is not a defense.
  4. Deciding what's worth investigating is a taste and priority call. A model ranks hypotheses by internal plausibility metrics; it doesn't know that your field's funders just shifted priorities, that a competing lab is six months from scooping a result, or that a "boring" hypothesis would actually resolve a decade-old controversy if someone finally ran the tedious experiment. That's a judgment about scarce attention and career risk that sits outside what the model is optimizing for.

New axioms

  1. Hypothesis overproduction outpaces the field's capacity to test anything. When generating 500 plausible hypotheses costs the same as generating 5, the bottleneck fully relocates to experimental throughput — and nobody has resolved how a lab should triage a firehose of AI-ranked candidates against a fixed budget of bench-hours. The failure mode is a lab drowning in "interesting" hypotheses it can never get to.
  2. Confidently-plausible-but-wrong hypotheses get funded and pursued at a larger scale than before. A hypothesis that pattern-matches well against the literature can still be mechanistically wrong, and AI ranking systems built on textual plausibility will systematically favor hypotheses that sound like existing successful papers — a citation-shaped prior, not a truth-shaped one. This risks a wave of well-ranked, well-funded, ultimately false leads, and nobody has built a reliable way to detect "sounds right because it's common in the literature" versus "sounds right because it's true."
  3. Novelty and priority claims get murkier when a model surfaces the same hypothesis to many labs at once. If several groups pull a similar AI-ranked hypothesis from the same broad training exposure, who gets credit, and was it really "discovered" independently? The field has no norm yet for what counts as an original contribution when the ideation step is a shared, non-scarce resource.
  4. Verifying that an AI's hypothesis ranking reflects real mechanism and not literature-frequency bias is itself a research problem with no settled method. Confirming a model's stated "confidence" or rank order tracks truth (not just training-data density) requires ground-truth validation nobody has fully solved — this is a fast-moving area and any claim about how well current systems do this should be treated as provisional.

Where it breaks

Labs treat "the model proposed it and ranked it highly" as if that were evidence of mechanistic plausibility, when the ranking is really a proxy for how well the hypothesis matches existing published patterns — collapsing the INVALID habit of trusting a well-read expert's gut ranking into an equally ungrounded trust in a model's ranking, while the genuinely NEW problem (hypotheses outpacing test capacity, and no way to separate citation-shaped plausibility from truth-shaped plausibility) goes unaddressed. The field replaced one unverified filter with another, faster one, and hasn't yet built the validation step that would make either trustworthy at scale.

Related axioms

Other axioms