No. 125 / 339
Do I still need to run real user interviews when synthetic users can simulate them in minutes?
The shift
LLMs make the synthesis and generation of plausible interview-shaped output abundant — a model can produce a coherent, well-articulated persona response to any question in seconds, at near-zero cost, at unlimited volume. What stays scarce is ground truth about what a real, specific human actually did, felt, or will pay for — a synthetic user has no data to report, only patterns to recombine.
The axioms
- Interviews exist to surface information the researcher doesn't already have. Rests on the interviewee holding facts, preferences, or context outside the researcher's model of the world.
- A user's answer is evidence about behavior, not just a plausible-sounding statement. Rests on the answer being anchored to a real, unrepeatable set of experiences — scarce because each person has lived only one life.
- Interviews are expensive to run, so teams ration them — small samples, tight scripts, careful recruiting. Rests on human time and attention being scarce.
- Synthesizing patterns across many past interviews and transcripts into themes takes analyst time. Rests on synthesis being slow when done by hand.
- Talking to users builds relationship and trust that later opens doors — for recruiting, for candor, for access. Rests on genuine rapport being something only formed between two real parties.
- Novel, unscripted follow-up questions ("wait, why did you say that?") depend on a listener who actually understood the specific answer just given, not a template. Rests on real-time judgment under ambiguity.
- Someone has to be accountable when a launch decision built on "we talked to users" turns out wrong. Rests on accountability requiring a party who can be blamed, fired, or who bears consequences.
Invalid axioms
- Interviews as the default first step to explore a hypothesis space or stress-test question wording. Generating plausible reactions, spotting confusing phrasing, or pressure-testing a script against many angles is now abundant — a model can role-play a dozen persona types on a draft script in minutes. The habit-trap: teams still burn a full recruiting-and-scheduling cycle just to learn "this question is ambiguous" or "this concept needs a one-line explainer" — information synthetic runs surface for free, before a real human's time gets spent on it.
- Manual thematic coding of large transcript archives to find recurring pain points. Synthesizing volume into themes is now cheap and fast. The habit-trap: research teams still staff for weeks of manual tagging on existing interview corpora instead of treating that layer as instantly automatable, saving the scarce human hours for what still needs a live person.
- Running interviews to rehearse the researcher's own framing before talking to anyone real. Practicing question order, checking for leading language, anticipating objections — a model can play adversarial or confused respondent instantly. The habit-trap: skipping this free rehearsal step and discovering script problems live, burning real participant goodwill on fixable mistakes.
Unchanged axioms
- A user's answer is evidence about behavior, not just a plausible statement. A synthetic user has no ground truth to draw on — it recombines what's already been written about people "like this," which means it reproduces the median assumption, not the surprising outlier that real interviews exist to catch. This is exactly the thing interviews are for, and it's the one thing the abundant side of AI cannot manufacture.
- Someone is accountable when a decision built on "we talked to users" goes wrong. A model can't be held responsible for a bad launch call, and "the synthetic panel said X" will not survive a postmortem the way "we interviewed twelve churned customers" does. Accountability requires a party that can answer for the claim, and that's still human.
- Trust and access that open doors for future recruiting, candor, and deeper follow-up. Rapport formed with a real person compounds — they refer other users, they tell you the thing they were embarrassed to say at first. A synthetic session produces no relationship and no future access; it's a dead end after the transcript.
- Judgment on the unscripted follow-up when an answer doesn't fit the pattern. A live researcher can catch a contradiction, sense hesitation, and probe the exact ambiguity that just appeared — that's improvisation against a specific, novel signal, not retrieval. Current models are fluent role-players but they're pattern-completing a persona, not noticing that this particular person just contradicted themselves in a way nobody has written about before.
New axioms
- Confidently wrong personas get treated as data because they arrive fast, well-written, and in bulk. When a synthetic panel returns 50 clean, articulate responses in the time one real interview takes, the sheer volume and fluency create an illusion of rigor — the open problem is building the discipline (and the org pressure) to keep treating synthetic output as a hypothesis generator, not a finding, when everything about its packaging says "finished analysis."
- Where does synthetic-user practice end and real-user validation legally/ethically have to begin? As synthetic research gets cheap enough to fully replace small-sample interviews in some teams' workflow, there's no settled answer for which decisions (pricing, safety-relevant UX, regulated claims) require a documented real-human evidence trail versus when synthetic pre-work is sufficient — this is still being negotiated inside product orgs, not resolved.
- Synthetic users trained on public discourse skew toward whoever writes the most online, silently erasing the users least represented in training data. The people most valuable to interview — because they're least legible from existing text (non-native speakers, low-literacy users, users in contexts underrepresented online) — are exactly the ones a synthetic panel will approximate worst, and there's no clean signal in the output telling you which persona responses are solid pattern-matches versus confident guesses about an undersampled group.
Where it breaks
Teams that swap real interviews for synthetic panels on early-stage concept testing (INVALID #1) will, without noticing, be making launch decisions off "confidently wrong personas treated as data" (NEW #1) — the synthetic step was meant to be disposable rehearsal, but because it's fast and well-written it quietly becomes the evidence base cited in the deck, right up until someone asks who's accountable for the miss (STILL HOLDS #2) and the answer is nobody, because no real user was ever in the room.
Related axioms
Research
What changes for scientific research with AI?
Research
Who's accountable when a product decision is made on synthetic-user data that turns out not to reflect real users?
Research
Is hypothesis generation still a scientist's job when AI systems can propose and rank novel hypotheses themselves?
Research
How does lab structure change when one PI plus AI agents can do the throughput that used to require five postdocs?
Research
How do we measure research team impact when "insights delivered" is no longer a scarce output?
Research
Should research ops still gatekeep access to real participants now that synthetic panels are the default first pass?
Other axioms
Legal
Who's liable when in-house counsel signs off on an AI-drafted contract that turns out wrong?
Engineering
What changes for DevOps with AI?
Hospitality
What changes for hospitality with AI?
Society
Does AI close the resource gap between small businesses and large competitors, or widen it, since large firms can deploy the same AI at far greater scale?
Industries
What changes for supply chain with AI?
Marketing
Is creative direction still scarce, or did AI just make "good enough" creative abundant and taste the only differentiator?