No. 123 / 339

What's the point of training PhD students on tasks an AI agent already does end to end?

The shift

AI makes abundant what a PhD used to be the only affordable way to buy: competent execution of a research task — running the literature review, writing the analysis code, drafting the lit-grounded first pass of a paper, debugging the pipeline. An agent with long context, tool use, and code execution now does the mechanical middle of research (search, synthesize, code, draft, iterate) at a speed and volume no student can match, and at near-zero marginal cost.

The axioms

  • The only way to learn to do research is to do the tasks a researcher does. (scarcity: skill was assumed inseparable from task-labor; you couldn't "practice" research without producing real research output)
  • A PhD is a multi-year apprenticeship priced in years of cheap, mistake-prone labor. (scarcity: expert supervision time was the bottleneck, so trainees paid for it in slow, imperfect output before earning autonomy)
  • Grinding through the slow, tedious version of a task is what builds the judgment to do it fast later. (scarcity: judgment was assumed to only form under the friction of doing the task yourself, at real speed, with real stakes)
  • The credential (the PhD) certifies that a person can be trusted to originate and stand behind research. (scarcity: trust and accountability can't be manufactured quickly — they require a track record)
  • Advisors can tell competent research from convincing-looking research because they've done the boring version themselves. (scarcity: the ability to spot a wrong-but-fluent result requires having built ground-truth intuition the hard way)
  • The point of a PhD is to produce novel knowledge, not to reproduce known procedure. (scarcity: genuinely new questions and judgment calls about what's worth asking were always the rare part; procedure was never the point, just the training vehicle)

Invalid axioms

  1. The only way to learn to do research is to do the tasks a researcher does. AI flipped the abundance of execution — running a regression, writing boilerplate code, summarizing fifty papers, drafting a section — from scarce-and-slow to instant-and-cheap. The habit-trap: programs still gate progression on hours spent manually producing these artifacts (write the code yourself, do the lit review by hand) as if the labor itself were the lesson, when the agent can produce a first pass and the actual lesson is now catching what it got wrong.
  2. A PhD is priced in years of cheap, mistake-prone labor for the advisor's lab. Labs used students partly as inexpensive execution capacity — someone to run the fiftieth variant of an experiment or clean a dataset. That labor is now abundant and near-free via agents. The habit-trap: funding models and lab throughput expectations still implicitly price a student's value by execution volume, which no longer differentiates them from a laptop.

Unchanged axioms

  1. The credential certifies someone can be trusted to originate and stand behind research. An agent can produce a paper-shaped artifact; it can't be an author who answers for a retraction, defends a result under adversarial questioning, or is accountable to a field's norms. Accountability is a relationship between a person and a community, not an output property — this doesn't move because the artifact got cheaper.
  2. Advisors can tell competent research from convincing-looking research because they've done the boring version themselves. This is the sharpest one, and it's getting more load-bearing, not less: agents raise the ceiling on producing fluent, plausible-looking research output, which raises the value of the specific skill of catching confidently-wrong results. That skill is currently built through exactly the friction the field wants to route around. If it erodes, verification capacity erodes with it — nobody left who's slow enough to have noticed the errors by hand.
  3. Judgment under novel, high-stakes ambiguity — what question is even worth asking — has no pattern to match. Agents are excellent at extending known methods to adjacent problems; they're weak at recognizing that a field's whole framing is wrong, or that a null result is the interesting finding, or that a reviewer's objection reveals a flaw nobody flagged. That call is still made by a person who's spent enough time in the weeds of a specific problem to have taste about it.

New axioms

  1. If agents generate the plausible first draft, what replaces slow labor as the forcing function for judgment? Grinding was never valuable for its own sake — it was the only available mechanism that happened to also build calibration. Abundance removes the mechanism without automatically replacing it. Nobody has a validated substitute yet: does reviewing and correcting an agent's output build the same intuition as producing the output yourself, slower and worse, a hundred times? This is genuinely open and moving fast — it hinges on model reliability, which is improving quarter over quarter.
  2. How does a field verify a PhD's competence when the artifacts they produce no longer prove it? A dissertation full of clean code and thorough lit review used to be weak evidence of competence because it was hard to fake at volume. An agent can now produce a competence-shaped artifact regardless of the student's actual skill. Evaluation hasn't caught up — most programs still grade the artifact, not the traceable judgment behind it.
  3. Who absorbs the cost when a lab runs on agent-verified-by-junior-researcher instead of junior-researcher-verified-by-senior-researcher? If students spend more time checking agent output than producing their own, the error-catching burden shifts to people with the least-developed ground-truth intuition — precisely because that intuition used to come from the labor now being skipped. This is a specific version of problem 1, sharp enough to name separately: it's a verification bottleneck stacked on top of a training bottleneck.

Where it breaks

Programs are already cutting the "produce it yourself" labor (INVALID 1) to save time — assigning agents to draft literature reviews and boilerplate analysis so students can "focus on higher-level thinking" faster. But the mechanism that built error-catching judgment (STILL HOLDS 2) was that exact labor. Removing it to move faster produces researchers who reach the point of independent judgment having never built the calibration to know when the agent is confidently wrong — which is the one moment in a research career when that calibration matters most.

A second collision: labs are also compressing the lab-throughput apprenticeship (INVALID 2), using freed-up junior time to run more agent-driven projects in parallel. But that scales problem 3 directly — more agent output, reviewed by people with less-formed judgment, at a volume no advisor can spot-check. The field is optimizing for throughput on the exact axis (verification capacity) it just made scarcer.

Related axioms

Other axioms