No. 116 / 339
Who's accountable when a product decision is made on synthetic-user data that turns out not to reflect real users?
The shift
Generating a plausible user — a synthetic persona's answers to a survey, an interview transcript, a usability reaction — is now near-zero cost and instant, where getting five real users into a room used to take a recruiter, an incentive budget, and a week. AI made the simulation of user opinion abundant; it did nothing to make that simulation true.
The axioms
- Research findings were trustworthy by default because producing them was expensive. Cost acted as a de facto quality filter: if someone spent the budget to run a study, the org treated the output as earned, not just plausible. Rests on scarcity of research capacity.
- The researcher who ran the study is accountable for its validity, because they chose the method, sample, and interpretation. Rests on scarcity of a visible, named human in the loop.
- A finding is credible in proportion to how many real humans it touched. Sample size was a proxy for truth because collecting real responses was slow and scarce.
- Product decisions get scrutinized in proportion to how expensive the input research was — cheap research got a skeptical read, expensive research got deference. Rests on cost as a stand-in for rigor.
- "The data reflects real users" was a checkable claim only after the fact, when the product shipped and usage data came in — so accountability was always somewhat retrospective and diffuse. Rests on scarcity of pre-launch ground truth.
Invalid axioms
- Cost as a quality filter for research. Expensive-to-produce research earned trust by default; AI made producing research-shaped output free, so cost no longer separates rigor from guesswork. The habit-trap: teams still wave a slide with "50 synthetic personas interviewed" as if the volume alone signals diligence, when volume is now the cheapest part of the process.
- Sample size as a credibility proxy. When real responses were scarce, more of them meant more truth. Synthetic samples can be generated in the thousands at no marginal cost, so "n=1000" from an LLM proves nothing about representativeness — it proves the loop ran a thousand times. The habit-trap: reports that quote synthetic sample sizes as if they carry the same evidentiary weight as recruited-user sample sizes.
- Research as a scarce, gating step before a decision. Because real research took weeks, decisions waited on it, and that wait itself forced deliberation. Now a team can generate a "user reaction" in minutes at any point in a meeting, which removes the forced pause — the gate is gone, not because the question got easier, but because the simulated answer got faster than the real one.
Unchanged axioms
- Someone specific has to own the call. A model can generate a synthetic user's opinion, but it cannot be held accountable when the decision built on it turns out wrong — there's no one to fire, no judgment to revoke, no reputation at stake for the LLM. Accountability still has to land on a named human: whoever decided synthetic data was sufficient evidence for this specific decision, at this specific stakes level. This doesn't move — AI has no capacity to hold it.
- Ground-truth verification that synthetic output matches real behavior. Nothing about generating a synthetic persona confirms it resembles the target population — that check requires actual users, actual usage data, or actual field validation. LLMs simulate plausible responses; they don't verify accuracy against reality. Whether synthetic data is "close enough" for a given decision is a judgment call, not something the model can self-certify.
- Judgment about when synthetic data is an acceptable substitute versus a false economy. Deciding "this decision is low-stakes enough that a synthetic proxy is fine" versus "this decision needs real users because getting it wrong is expensive to reverse" is a call under ambiguity — no pattern to match, no way to outsource it. This is a taste-and-stakes judgment, and it still sits with the human who greenlights the research plan, not the tool that executes it.
- Trust between the product team and whoever consumes its findings. A stakeholder signs off on a roadmap because they trust the team's research process, not because they personally re-derive it. That trust relationship — and the standing to say "I vouch for this" — is a human commitment a model cannot make on anyone's behalf.
New axioms
- Synthetic data can look and read exactly like real research, with no visible seam. A transcript from a synthetic-user interview and one from a real recruited participant are formatted identically, which means downstream reviewers — execs, other PMs, auditors — have no way to tell which is which unless the team explicitly labels it. What provenance-tagging or disclosure norm makes the origin of a finding visible at the point of decision, not buried in a methods appendix nobody reads?
- The failure is discovered only after the product ships, by which point the "decision" has fanned out into roadmap commitments, code, and marketing. Because synthetic research is cheap and fast, teams can now run more of it and make more decisions off it per unit time — which means more decisions are exposed to this failure mode simultaneously, and the retrospective accountability process (which was already slow) hasn't sped up to match. How does a team build a fast-enough feedback loop to catch a bad synthetic-data decision before it compounds into three sprints of built features?
- Synthetic personas trained on average patterns systematically miss the users who matter most for a given decision — edge cases, non-English speakers, accessibility needs, adversarial or bad-faith users, the exact minority behavior a product decision often hinges on. Volume of synthetic output doesn't fix this; it can mask it, because a thousand synthetic responses that all miss the same blind spot look like consensus. Who's responsible for defining which decisions are structurally unsafe to make on synthetic data at all, given the model's known skew?
- Accountability structures (sign-off chains, research review boards, "who approved this" logs) were built for a world where research was rare and expensive, and haven't been redesigned for a world where anyone can generate a study in an afternoon. Without a process update, accountability defaults to whoever happened to run the prompt — usually the most junior person in the room — rather than whoever should own a decision at that stakes level.
Where it breaks
A PM runs synthetic-user interviews in an afternoon (INVALID #1, #3 — the cost-and-time gate that used to force scrutiny is gone) and ships a roadmap decision off them. Three months later usage data shows the synthetic personas missed a real segment (NEW #3). Nobody can reconstruct who signed off on treating synthetic data as sufficient for that decision, because the research review process still assumes research is rare enough that everyone remembers who ran it (NEW #4) — the volume that made the mistake possible also erased the paper trail that would assign blame for it.
A second collision: leadership sees a synthetic research deck formatted identically to a real one (NEW #1) and defers to it the way they used to defer to expensive research (the now-dead cost-as-quality-filter, INVALID #1) — the visual and structural cues that once correlated with rigor now correlate with nothing, but the deference habit hasn't updated.
Related axioms
Research
What changes for scientific research with AI?
Research
Is hypothesis generation still a scientist's job when AI systems can propose and rank novel hypotheses themselves?
Research
How does lab structure change when one PI plus AI agents can do the throughput that used to require five postdocs?
Research
How do we measure research team impact when "insights delivered" is no longer a scarce output?
Research
Should research ops still gatekeep access to real participants now that synthetic panels are the default first pass?
Research
Is the paper still the right unit of scientific output when AI can generate them faster than humans can read them?
Other axioms
Society
What changes for social work with AI?
Industries
What changes for upstream oil and gas with AI?
Cybersecurity
Who's liable for a breach an AI security agent missed or misclassified?
Marketing
Do we still need product marketing to translate features into positioning if AI drafts messaging from changelogs?
Healthcare
If an AI scribe drafts the progress note from a session, who owns the medical-necessity language and the liability when it's wrong?
HR
Should performance reviews still be written manually when AI can draft them from a manager's notes and work history?