No. 119 / 339

How do we measure research team impact when "insights delivered" is no longer a scarce output?

The shift

AI collapses the cost of producing a plausible-sounding insight — a theme summary, a persona, a "customers want X" writeup — to near zero and seconds, via synthesis of large volumes of qualitative and behavioral data. Anyone with access to a transcript pile or a dashboard can now generate what used to require a trained researcher's time, so volume and speed of "insights delivered" stop signaling that a research team did anything only it could do.

The axioms

  • Insight production is scarce, expensive, trained labor — rests on synthesis requiring a researcher's time and judgment.
  • Volume of insights delivered tracks team value — rests on insight production being the bottleneck, so more output means more value created.
  • A researcher's name on a finding is what makes stakeholders trust it — rests on scarce visible rigor (the credentialed process) as the credibility signal.
  • Research team headcount should scale with the number of questions the business wants answered — rests on each answer costing a fixed amount of scarce analyst time.
  • The research report or readout is the unit that proves work happened — rests on the deliverable being expensive to produce, so its existence is evidence of effort.
  • Research influence is measured by how many findings get cited in a roadmap doc — rests on citation being a rare, deliberate act signaling someone actually engaged with real work.

Invalid axioms

  1. Volume of insights delivered tracks team value. AI makes synthesis and first-draft "insight" generation abundant — a PM, a founder, or an AI agent pointed at the same data can produce comparable-looking output in minutes. The habit-trap: teams still report headline metrics like "insights shipped this quarter" or "studies completed," rewarding throughput that no longer differentiates a research team from a well-prompted intern with an LLM.
  2. A researcher's name on a finding is what makes stakeholders trust it. Credibility used to ride on the visible labor of the process — weeks of fieldwork nobody could fake. Now that AI can generate a confident, well-formatted, plausible-sounding finding just as easily whether or not it's grounded in real signal, attaching a researcher's name to output no longer functions as a trust signal on its own. The trap: research orgs keep branding deliverables ("Research says...") as if the label still carries the old weight.
  3. The research report is the unit that proves work happened. When a polished readout can be drafted by AI from a stack of transcripts in an afternoon, the artifact stops being evidence of researcher effort. Teams that still measure themselves by reports shipped or decks produced are counting an output that costs nothing to fake.
  4. Headcount should scale with the number of questions the business wants answered. Question-answering throughput is exactly what AI synthesis makes cheap. Sizing a research team to the volume of asks — rather than to the number of high-stakes, ambiguous calls that need real verification — staffs for a bottleneck that no longer exists.

Unchanged axioms

  1. Someone must be accountable when a decision built on research turns out wrong. AI can generate a finding; it cannot be held answerable when a launch fails because the finding was confidently wrong. Impact measurement that matters has to trace back to a person who owns the call, not to an insight-count.
  2. Getting to a genuinely novel, unarticulated need still requires talking to real, specific humans under real stakes. Synthesis can summarize what's already been said; it can't surface what no one has said yet. A research team's real differentiator — finding the thing nobody knew to ask about — still rests on scarce, non-substitutable human contact, and that work is measurable by outcome (did it change a decision) even though the underlying labor is undervalued by output-counting metrics.
  3. Deciding which questions are worth answering is a judgment call, not a production task. AI increases the supply of answers; it does nothing to increase the supply of good judgment about which ambiguous, high-stakes question deserves scarce verification effort. A research team's impact is disproportionately concentrated in a small number of correctly chosen bets, and that selection judgment is still the scarce skill worth measuring.
  4. Stakeholder trust, once earned, still tracks a track record of being right, not speed or volume. A team that's been wrong before doesn't get more credibility by producing more findings faster — trust is relational and accumulates through being verifiably correct over time, which AI-generated volume cannot substitute for.

New axioms

  1. There's no metric yet for "insight that was actually verified against reality" versus "insight that merely sounded right." Output volume used to correlate loosely with rigor because production was slow and expensive; now the correlation is gone, and no standard measure fills the gap — teams need something like a verification rate or decision-accuracy track record, and it doesn't exist as a norm yet.
  2. Decision quality downstream of research is now the only signal left that isn't gameable, but it's slow, noisy, and hard to attribute. A launch's success six months later reflects many factors besides the research that informed it; building an impact metric around downstream outcomes means solving attribution and lag, which nobody on a research team has historically had to do because the input metric (insights produced) was good enough to prove the team existed.
  3. AI-abundant insight production creates a flood that overwhelms whoever is supposed to sanity-check it, and no one owns that sanity-check function. If everyone in the org can generate "research-flavored" output, someone has to triage what's grounded from what's plausible noise before it hits a roadmap — that gatekeeping job is newly necessary and currently unassigned, and it's the actual thing worth measuring a research team's impact against, since catching bad synthesis before it ships is now more valuable than producing more synthesis.
  4. Research team value is at risk of being invisible precisely because its best work (calling the right high-stakes question, catching a wrong synthetic finding before launch) doesn't show up as an artifact. The old metric was visible (reports, studies, insights); the surviving value is a judgment call that, when done well, looks like nothing happened — this makes a research team's real impact harder to defend in a budget conversation, not easier, even though the value increased.

Where it breaks

Leadership still reports and budgets research teams on "insights delivered" (invalid: volume tracks value) while the actual differentiating work has shifted to catching wrong AI-generated synthesis before it reaches a roadmap (new: nobody owns the verification/triage function) — so the team gets measured on the exact output that's now cheapest and least differentiating, while the scarce judgment work that justifies the team's existence goes unmeasured and undefended.

A second collision: research orgs keep branding findings with a researcher's name as the credibility signal (invalid: the name used to prove costly rigor happened) at the same moment stakeholders can't tell a verified finding from a plausible AI-generated one without asking (new: no metric exists for verified-versus-plausible) — so the exact mechanism research teams rely on to be trusted is the one AI has quietly hollowed out, and nobody has replaced it with a check that actually distinguishes the two.

Related axioms

Other axioms