No. 140 / 339

Does "showing your work" still signal trust when the work itself is trivially reproducible?

The shift

Producing a plausible, polished work-trail — a documented process, a worked derivation, a design rationale, a paper trail of visible effort — goes from costly to near-free. This is the "generate plausible first-drafts" and "translate between levels of expertise" capability pointed at the artifact of effort itself: an LLM will reconstruct a convincing reasoning chain, a tidy commit history, or a step-by-step justification for a conclusion it never actually reasoned through. The signal that shown work used to carry — "faking this was nearly as expensive as doing it, so its presence is evidence of real effort" — is what breaks.

The axioms

  • A visible process trail is evidence of genuine effort — it rests on the cost of fabricating a convincing trail being close to the cost of doing the work for real, so the trail's existence certified the work.
  • A worked-out rationale demonstrates understanding — it rests on articulation being scarce: only someone who actually understood could produce a coherent, defensible explanation.
  • Grading, reviewing, and trusting the artifact (the shown work) is a fair proxy for grading the competence behind it — rests on artifact quality and underlying competence being tightly coupled, because producing the artifact required the competence.
  • A verifiable, correct result is trustworthy — rests on ground truth being checkable independent of who or what produced it.
  • Staking your name on a claim signals trust — rests on reputation being costly to build and costly to lose, so putting it on the line is a credible bond.
  • Accountability for the outcome signals trust — rests on there being a party who bears real consequences when the result is wrong.

Invalid axioms

  1. A visible process trail is evidence of genuine effort. The trail was trusted because faking it cost about as much as doing the work. That coupling is severed: a fluent rationale, a plausible methodology section, a clean-looking audit trail, or a reconstructed chain of reasoning is now generated in seconds regardless of whether any real work happened. Habit-trap: teachers, reviewers, hiring panels, and managers still ask people to "show your work" and treat the presence of shown work as reassurance — grading the artifact of effort rather than probing the thing effort was supposed to produce.
  2. A worked-out rationale demonstrates understanding. Articulation used to be a reliable tell for comprehension because only understanding could generate a coherent defense. LLMs make coherent post-hoc justification abundant and decoupled from understanding — a confident, well-structured explanation is exactly what the model produces most easily, including for conclusions that are wrong. Habit-trap: interviews, design reviews, and oral defenses that probe "walk me through your reasoning" still read fluency as a proxy for depth, when fluency is now the cheapest thing in the room.
  3. Grading/reviewing/trusting the shown work is a fair proxy for the competence behind it. Artifact and competence have come uncoupled: the artifact can be produced without the competence, and increasingly the competence (e.g. someone who genuinely understands but uses AI to write it up) produces an artifact indistinguishable from a fabricated one. Habit-trap: institutions built to score the visible output — essays, take-home assignments, process docs, documented decision rationales — keep optimizing the thing that no longer discriminates between the competent and the confabulating.

Unchanged axioms

  1. A verifiable, correct result is trustworthy. Cheap-to-fake process doesn't make ground truth cheap to fake. A proof that machine-checks, code that passes an adversarial test suite, a prediction that comes true, a number that reconciles against an independent source — these still cost what they cost, and the fabricated version fails the check. The move is from trusting the trail to trusting the outcome, and where the outcome is independently verifiable, trust survives intact. The load-bearing caveat: this only holds when a real, hard-to-game ground truth exists and someone actually runs the check. For claims with no cheap oracle — strategy, taste, "is this the right architecture" — there is no result to verify, and this bucket offers no shelter.
  2. Staking your name on a claim still signals trust. Reputation stays costly to build and costly to lose, and AI doesn't let you fake having something to lose. A named person who will be blamed if the work is wrong is putting up a bond the model can't post. What's changed is scope, not validity: the signal now attaches to the person who vouches, not to the artifact they attached their name to — the polished trail no longer adds to the bond, only the staking does.
  3. Accountability for the outcome still signals trust. Someone answerable when the result fails — who bears the cost, the liability, the lost client — is a scarcity AI doesn't touch, because a model can't be the answerable party. This is the durable floor: when process, rationale, and even shown results can all be manufactured, "who eats the loss if this is wrong" remains a real and un-fakeable signal.

New axioms

  1. When the process trail is free to fake, trust has to move to outcomes, provenance, or reputation — and we don't have cheap mechanisms for the latter two yet. Verifiable outcomes cover only the slice of work with a cheap ground-truth oracle. For everything else, the fallback is provenance (who/what actually produced this, through what chain) and reputation (whose name is on it) — but there's no cheap, trustworthy provenance layer today. Cryptographic signing, content credentials, and audit-logging exist but aren't ambient, aren't universally adopted, and mostly prove "a tool touched this," not "a competent human stood behind it." This is moving fast: provenance/attestation standards and hardware-backed content credentials are an active build-out, so how NEW this stays is a near-term open call.
  2. We're over-trusting shown work in exactly the domains where no ground-truth check exists. The verifiable-outcome shelter creates a false sense of coverage: fields that can run the check (code, math, forecasting) adapt, while fields that can't (strategy memos, policy rationales, qualitative analysis, most managerial and design judgment) still lean on shown reasoning as their only trust signal — precisely the signal that just died. The open problem is finding a substitute for "show your work" where there's nothing to verify the work against.
  3. The trail is now negative evidence in a way it never was. A suspiciously clean, comprehensive, instantly-produced rationale used to read as diligence; it now reads as a possible tell for automation. This inverts the incentive — visible effort can lower trust rather than raise it — and we have no shared norms for what a credible work-trail looks like in a world where polish is free. Roughness, friction, and the marks of real struggle may become the scarce signal, but there's no established way to demand or verify them.

Where it breaks

Institutions still grade and trust the shown-work artifact (INVALID #1, #3) in exactly the domains where no ground-truth oracle exists (NEW #2) — education's essays and take-homes, the design review's "walk me through your thinking," the consultant's methodology section. The verifiable-result shelter (STILL HOLDS #1) doesn't reach these, so the fields most dependent on shown work as their trust mechanism are the ones whose trust mechanism just stopped working, and they're responding by demanding more shown work — manufacturing more of the free, undiscriminating signal.

A second collision: the polished trail is turning into negative evidence (NEW #3) while reputation and accountability are the only bonds that still hold (STILL HOLDS #2, #3) — yet most systems still let people attach their name to AI-generated work-trails cheaply, with no provenance layer (NEW #1) to distinguish "I stood behind this" from "I signed off on output I didn't verify." The staking signal stays valid in principle but gets diluted in practice, because the act of vouching has been made as cheap to perform as the trail it's vouching for.

Related axioms

Other axioms