No. 122 / 339

Is peer review still meaningful when a meaningful share of reviews at major venues are already AI-generated?

The shift

Producing a fluent, structured, plausible-sounding review — complete with citations, framing against related work, and specific-sounding critique — goes from scarce (an expert's unpaid hours) to abundant, near-instant, and effectively free. Verifying whether that critique is actually correct, or whether a human engaged with the paper at all, stays hard and stays manual.

The axioms

  • A qualified expert reads and evaluates the submission themselves — rests on reviewer attention being genuinely, scarcely applied.
  • The review reflects the reviewer's own judgment and their name/reputation backs it — rests on accountability: someone stakeable is answerable for the verdict.
  • Reviewing takes real effort, so agreeing to review signals baseline commitment — rests on review labor being costly enough to filter out low-commitment reviewers.
  • A small pool of qualified people is available per paper, so venues ration reviewer slots — rests on expertise being scarce relative to submission volume.
  • A detailed, articulate review is evidence of deep engagement — rests on fluent critique being expensive to produce, so its presence signaled effort.
  • Peer review catches errors, fraud, and weak methodology through genuine comprehension — rests on ground-truth verification requiring real understanding.
  • The volume a venue can fairly evaluate is bounded by willing, competent reviewers — rests on reviewer supply being scarce and slow to expand.

Invalid axioms

  1. A detailed, articulate review is evidence of deep engagement. Fluency and structure used to correlate with effort because writing a coherent multi-paragraph critique took real time. AI collapses that cost to zero, so length and polish stop being a signal of anything. The habit-trap: editors and program chairs still read a well-organized review as evidence of care, and reviewers who paste a model's output get credit for engagement they didn't provide.

  2. Venues ration reviewer slots because expertise is scarce relative to submission volume. A reviewer with a model can now produce competent-sounding coverage of a paper's methods section, related work, and framing in minutes, so venues can nominally cover more submissions per reviewer without adding reviewer-hours. The habit-trap: program committees still size reviewer pools and set review deadlines as if the constraint were expert time, when the actual bottleneck has moved to whether anyone is checking the reviews themselves.

  3. Reviewing takes real effort, so agreeing to review signals baseline commitment. The norm that a review's existence implies the reviewer invested comparable effort to the author no longer holds — a reviewer can generate a full-length critique with a fraction of the effort a submission took to write. The habit-trap: acceptance/rejection decisions still weight review length and specificity as a proxy for reviewer seriousness.

Unchanged axioms

  1. The review must reflect the reviewer's own judgment, backed by their name. Accountability didn't get cheaper. A model can produce a critique, but it can't be the answerable party when a review is wrong, biased, or missed fraud — that's still a named human's reputation and (in some venues) career incentive on the line. The scandal isn't "AI helped write a review," it's a human submitting AI output under their own name without owning its correctness.

  2. Peer review catches errors and weak methodology through genuine comprehension of the specific paper. Current models are good at generating plausible-sounding methodological critique — pattern-matched against thousands of prior reviews — but they're unreliable at verifying whether a novel result's central claim actually holds, spotting fabricated data, or catching a subtle flaw outside the training distribution. Confidently wrong is the default failure mode, and it's exactly as dangerous in a review as in the paper it's reviewing. This is worth flagging as unstable: as models get more agentic and better at tool use (running code, checking datasets, cross-referencing citations live), the verification gap narrows — but it hasn't closed yet.

  3. Judgment on novel, high-stakes calls — is this result actually important, does it deserve to be in this venue — requires a human weighing context the paper doesn't state. Deciding whether a marginal-but-technically-sound result matters, or whether a flashy result is being oversold, is a taste/judgment call, not a pattern-match against prior reviews.

New axioms

  1. Nobody can currently verify, at scale, whether a review was AI-generated, human-written, or some blend — and it's not clear a venue can build that detector faster than models improve at evading it. Review integrity now has the same provenance problem the papers themselves have.

  2. If a meaningful share of reviews are AI-generated, review quality stops correlating with reviewer scarcity, and venues have no replacement signal for "was this actually vetted." The old proxies (review length, specificity, turnaround time) are exactly the things AI is best at faking, so the field needs a new signal for genuine engagement that doesn't just reward fluency.

  3. Authors can now also run their own paper through the same models reviewers use before submission, so the loop of AI-critiquing-AI-refined-text risks converging on a local optimum of "passes model-style scrutiny" rather than "is true or important." This is a feedback loop that didn't exist when critique was expensive.

Where it breaks

Venues still ration reviewer slots as if expertise were the bottleneck (INVALID #2) while having no mechanism to verify whether the reviews filling those slots reflect real comprehension or AI-generated pattern-matching (NEW #1) — a program committee can hit its reviewer-per-paper quota and still have no idea whether the paper was actually checked.

Review length and specificity are still read as evidence of reviewer effort (INVALID #1 and #3) at the exact moment those are the cheapest things for a model to fake convincingly (NEW #2) — the signal the field is defending is the one that costs nothing to counterfeit.

Related axioms

Other axioms