No. 161 / 339

If AI designs the experiment, who owns the wet-lab execution and the reproducibility of the result?

The shift

Experimental design — the protocol, the controls, the statistical plan, the reagent list, the step-by-step method — goes from scarce, trained-scientist output to abundant, near-instant, near-free. What doesn't move: getting the physical run to actually work at the bench, and being able to reproduce that specific run. AI can specify the experiment in full; it still can't pipette, notice the plate looks wrong, or vouch that the data on the screen came from clean hands.

The axioms

  • The intellectual value of an experiment lives in its design; execution is downstream labor that a trained pair of hands carries out.
  • Skilled physical execution at the bench is scarce, and reliable data depends on it — the same protocol in two people's hands does not give the same result.
  • A result is reproducible because a named human ran it, recorded what actually happened, and can vouch that it's real rather than an artifact.
  • Catching contamination, edge cases, and artifacts requires trained physical intuition built up over years at the bench — the "that doesn't look right" reflex.
  • When a run fails or won't replicate, it is owned by the person who ran it; they debug it, and their name is on whether it's trustworthy.
  • Operating physical lab equipment requires certified human competency — someone signed off as trained to run the instrument, handle the hazard, and follow the protocol.
  • Reproducibility is a property that gets established after the fact, by another human repeating the work, because repeating it was expensive and rare.

Invalid axioms

  1. The intellectual value of an experiment lives in its design. Designing the protocol — controls, power calculation, reagent selection, step sequence, statistical plan — is now something a model produces in seconds, grounded in the literature and often better-specified than a rushed grad student's. The design was treated as the scarce, creditable contribution and execution as mere hands. The habit-trap: labs still credit, hire, and promote around who designs the study, and treat the bench work as junior labor to be delegated and forgotten — exactly inverting where the remaining scarcity now sits.

Unchanged axioms

  1. Skilled physical execution at the bench is scarce, and reliable data depends on it. A perfect protocol run by unsteady or untrained hands gives unreliable data; the same document in two labs still gives two results. AI-designed does not mean AI-executed — the throughput ceiling on pipetting, culturing, staining, and handling live systems has barely moved relative to how fast the design side collapsed. Lab robotics automates a slice of this, but only for assays already engineered for it, and someone still loads, calibrates, and babysits the robot.
  2. Someone accountable must vouch that a result is real and not an artifact. A model can output a protocol and, downstream, an analysis of the resulting data; it cannot stand behind the claim that the data came from an uncontaminated run done as written. Contamination, mislabeled tubes, a degraded reagent, a miscalibrated instrument — the accountable human who says "I ran this, it's clean, it's real" is still the verification layer, and can't be a model.
  3. Catching contamination, edge cases, and artifacts requires trained physical intuition. The reflex that a culture looks off, a band is in the wrong place, or a reading is too clean to be true is built at the bench and applied in the physical moment. A model reasoning over a description of the run is not in the room to see the plate.
  4. When a run fails or won't replicate, debugging it is physical, iterative, and human. Chasing down why a replication failed — was it the lot of antibody, the ambient temperature, the technique, the water — is troubleshooting in the physical world, not a token-generation problem. AI can suggest candidate causes; someone still has to go rule them out at the bench.

New axioms

  1. No validated framework exists for reproducibility of AI-designed wet-lab experiments, so error propagates silently. When the design is abundant and machine-generated, a flawed control, a subtly wrong concentration, or an inappropriate statistical plan can be replicated across many labs at once before anyone catches it — the error is now upstream of every execution instead of contained in one lab's judgment. There is no accepted standard for validating an AI design before hands touch it, and no framework for tracing an irreproducible result back to a design flaw versus an execution flaw.
  2. Nobody owns a failed or irreproducible AI-designed run. When the design came from a model and the run came from a technician following it, and the two disagree, accountability splits with no owner: the bench worker says they ran it as written, the design has no author who can be liable, and the PI approved a protocol they didn't write. The credit-and-fault chain that used to run from a human designer to a human executor now has a gap in the middle where the design came from.
  3. Automating the hands where no competency standard exists. As AI drives physical lab equipment directly — liquid handlers, automated culture systems, closed-loop "self-driving" labs — there is no competency standard for what an AI agent is certified to operate, what hazards it may handle unsupervised, or who is qualified to sign off that it ran safely and correctly. The certification that gates a human at the instrument has no equivalent for the agent now reaching for the same instrument.
  4. When designs are abundant, the bottleneck moves to bench capacity and no one has re-resourced it. A lab can now generate more validated-looking protocols than it can physically run or reproduce. The scarce resource becomes bench time and the accountable hands to execute and verify — and nothing in how labs are staffed or funded has shifted to match, because the design side is what still gets the budget and the headcount.

Where it breaks

"The value of an experiment is in its design" (invalid) collides with "no one owns a failed or irreproducible AI-designed run" (new): labs are eagerly adopting AI to generate protocols because design was always the prized, creditable work — while the execution and the vouching-for-the-result, now the actually scarce parts, are pushed onto technicians who didn't author the design and can't be held responsible for its flaws. The result that won't replicate has an author for its idea and an operator for its hands, and no one who owns both.

A second collision: "reproducibility gets established after the fact by a human repeating the work" (a holdover assumption, cheap when designs were scarce) meets "AI-designed experiments propagate error across many labs at once" (new). The field's reproducibility check was built to catch one lab's mistake at one lab's pace; when a single flawed machine-generated design seeds identical runs everywhere simultaneously, after-the-fact replication is checking downstream of the error instead of at it — and this is the call most likely to shift fast as self-driving labs and agentic equipment control mature, so it should be revisited against capability rather than treated as settled.

Related axioms

Other axioms