No. 117 / 339
Is hypothesis generation still a scientist's job when AI systems can propose and rank novel hypotheses themselves?
The shift
AI makes breadth of combinatorial search abundant — pattern-matching across a field's entire literature to propose large numbers of novel, mechanistically plausible hypotheses in minutes, and ranking them by fit to existing data. What stays scarce is knowing which of those hypotheses is worth the months of wet-lab or field time to actually test, and whether the ranking reflects real biological or physical plausibility versus literature-frequency bias.
The axioms
- A scientist's expertise is measured partly by the breadth of prior work they can hold in mind while generating ideas — this rested on human recall and reading speed being scarce.
- Novel hypotheses come from a trained expert noticing an unusual connection across disparate literatures — this rested on cross-domain synthesis being slow and rare.
- Generating many candidate hypotheses is costly, so scientists generate few and commit early — this rested on ideation itself being expensive.
- A hypothesis is only as good as the experiment that can falsify it, and someone has to design and run that experiment — this rests on physical action and lab access being scarce.
- Ranking hypotheses by plausibility requires judgment about mechanism, prior probability, and what's actually testable given real-world constraints (funding, ethics, time) — this rests on contextual judgment, not just pattern-fit.
- The person who proposes a hypothesis is accountable for the resources spent testing it and the claims made from the result — this rests on accountability being tied to a name.
- Deciding which question is worth asking at all — which gap in a field matters — is a taste and priority call, not a search problem — this rests on scarce judgment about what's worth doing.
Invalid axioms
- Breadth of recall is a proxy for scientific insight. AI now reads and cross-references the full literature of a field (and adjacent fields) faster and more completely than any human can. The habit-trap: labs still treat the "well-read" postdoc or the person who "knows everyone's work" as the ideation bottleneck, and route hypothesis generation through them, when a model can already surface the same connections faster and generate more candidates than that person would in a week.
- Generating many candidate hypotheses is expensive, so you generate a handful and commit. AI collapses the cost of producing dozens or hundreds of plausible hypotheses to near zero. The habit-trap: grant proposals and thesis committees still expect a single early-committed hypothesis with sunk reasoning behind it, rewarding the appearance of foresight over what's now cheap — actually generating and comparing a wide candidate set before committing lab time.
- Cross-domain synthesis (spotting a mechanism from field A that explains a puzzle in field B) is a rare individual talent. This was scarce because no one person reads deeply across enough fields. Models with broad training now do this by default. The habit-trap: institutions still prize and specially fund "interdisciplinary" hires as the mechanism for cross-pollination, when the pollination itself is now a commodity — the scarce part has moved to knowing which cross-domain suggestion is real versus superficial pattern-matching.
Unchanged axioms
- Someone has to design and run the experiment that falsifies the hypothesis. AI can propose a mechanism; it cannot pipette, culture cells, run a particle detector, or recruit trial participants. Hypothesis generation and hypothesis testing are decoupled by this shift, but testing remains physical-world-scarce, so the bottleneck in most fields simply moves downstream rather than disappearing.
- Ranking hypotheses by real-world testability and cost is a judgment call, not a scoring function. A model's ranking is trained on textual plausibility and citation patterns — it doesn't know your lab's equipment, your funding cycle, or that a particular reagent has an 18-month backorder. Deciding which of 200 AI-generated hypotheses is worth six months of grad student time is a judgment about scarce resources the model doesn't have visibility into.
- Accountability for a claim stays with a named scientist. If a hypothesis leads to a retraction, a wasted grant, or a harmful clinical recommendation, a person answers for it — a co-author, a PI, a reviewer. AI proposing the hypothesis doesn't create a locus of accountability; it just means the human who selected and pursued it owns the consequences more fully, since "the model suggested it" is not a defense.
- Deciding what's worth investigating is a taste and priority call. A model ranks hypotheses by internal plausibility metrics; it doesn't know that your field's funders just shifted priorities, that a competing lab is six months from scooping a result, or that a "boring" hypothesis would actually resolve a decade-old controversy if someone finally ran the tedious experiment. That's a judgment about scarce attention and career risk that sits outside what the model is optimizing for.
New axioms
- Hypothesis overproduction outpaces the field's capacity to test anything. When generating 500 plausible hypotheses costs the same as generating 5, the bottleneck fully relocates to experimental throughput — and nobody has resolved how a lab should triage a firehose of AI-ranked candidates against a fixed budget of bench-hours. The failure mode is a lab drowning in "interesting" hypotheses it can never get to.
- Confidently-plausible-but-wrong hypotheses get funded and pursued at a larger scale than before. A hypothesis that pattern-matches well against the literature can still be mechanistically wrong, and AI ranking systems built on textual plausibility will systematically favor hypotheses that sound like existing successful papers — a citation-shaped prior, not a truth-shaped one. This risks a wave of well-ranked, well-funded, ultimately false leads, and nobody has built a reliable way to detect "sounds right because it's common in the literature" versus "sounds right because it's true."
- Novelty and priority claims get murkier when a model surfaces the same hypothesis to many labs at once. If several groups pull a similar AI-ranked hypothesis from the same broad training exposure, who gets credit, and was it really "discovered" independently? The field has no norm yet for what counts as an original contribution when the ideation step is a shared, non-scarce resource.
- Verifying that an AI's hypothesis ranking reflects real mechanism and not literature-frequency bias is itself a research problem with no settled method. Confirming a model's stated "confidence" or rank order tracks truth (not just training-data density) requires ground-truth validation nobody has fully solved — this is a fast-moving area and any claim about how well current systems do this should be treated as provisional.
Where it breaks
Labs treat "the model proposed it and ranked it highly" as if that were evidence of mechanistic plausibility, when the ranking is really a proxy for how well the hypothesis matches existing published patterns — collapsing the INVALID habit of trusting a well-read expert's gut ranking into an equally ungrounded trust in a model's ranking, while the genuinely NEW problem (hypotheses outpacing test capacity, and no way to separate citation-shaped plausibility from truth-shaped plausibility) goes unaddressed. The field replaced one unverified filter with another, faster one, and hasn't yet built the validation step that would make either trustworthy at scale.
Related axioms
Research
What changes for scientific research with AI?
Research
Who's accountable when a product decision is made on synthetic-user data that turns out not to reflect real users?
Research
How does lab structure change when one PI plus AI agents can do the throughput that used to require five postdocs?
Research
How do we measure research team impact when "insights delivered" is no longer a scarce output?
Research
Should research ops still gatekeep access to real participants now that synthetic panels are the default first pass?
Research
Is the paper still the right unit of scientific output when AI can generate them faster than humans can read them?
Other axioms
Construction
How does construction project management change when AI coordinates subcontractor scheduling and logistics directly?
Management
Do we still need the analyst who builds slides if the AI builds better slides faster?
Media
When AI can master a track to "radio-ready" in one click, what stays scarce about the audio engineer's ear?
Industries
What changes for plumbers with AI?
HR
What changes for HR with AI?
Management
What changes for management consulting with AI?