No. 6 / 339

Who's liable when an AI-generated structural model passes every check but fails in the field?

The shift

Generating a plausible, code-compliant structural model — geometry, load paths, member sizing, checked against every codified rule — goes from scarce senior-engineer hours to abundant and near-instant. What doesn't get cheaper is verifying the model's assumptions against ground truth the code never encoded, and identifying who's answerable when a rule-passing model still fails.

The axioms

  • The engineer of record's stamp is the single point of accountability, because liability for public safety has to land on someone and can't be distributed.
  • Code-check compliance is a workable proxy for real-world safety, because exhaustively testing every load case and site condition is too expensive — codified rules substitute for it.
  • Professional judgment catches what the checks don't, because codes can't anticipate every material quirk or site condition; a scarce expert fills the gap.
  • Peer review and independent checking are worth their cost because building the model itself was expensive — check budgets are set as a fraction of build time, on the assumption both are scarce.
  • Insurance prices risk by tracing it to whoever exercised discretionary judgment, because liability follows the party whose decisions could plausibly have caused the failure.
  • Design tools are non-agentic — software calculates what a human directs, so the human is presumed author of every assumption baked into the model.

Invalid axioms

  1. Peer-review hours should scale with how long the model took to build. That budgeting logic assumed model production was the expensive, slow step and review was cheap by comparison. AI collapses production time toward zero while leaving verification just as hard — so a review budget still sized off build-time under-resources the one step that didn't get cheaper. Habit-trap: firms still quote and staff checking as a percentage of design fee, not as its own fixed cost driven by the model's actual novelty or risk.
  2. Code-check compliance signals the model is trustworthy. Passing every codified check was a reasonable stand-in for safety when models were hand-built by someone who'd also exercised judgment on everything the code didn't cover. When an AI generates the model, it optimizes to pass the checks, not to encode the judgment that made passing checks meaningful in the first place. The habit-trap: treating "it passed" as roughly the same signal it used to be, when it's now a weaker, gameable proxy — abundant plausible-and-compliant output, not abundant correctness.
  3. The human author of the model is presumably the source of its assumptions. Tools were non-agentic; whoever set the parameters owned them. Once an AI proposes load paths, boundary conditions, or member selections a human didn't originate, "the engineer authored every assumption" stops being literally true, even though the stamp still says they did. Habit-trap: sign-off workflows unchanged from the hand-calculation era, as if speed of drafting were still the bottleneck being managed.

Unchanged axioms

  1. The engineer of record is legally answerable, not the tool. Accountability doesn't transfer to software no matter how the model was produced — this is a licensing and legal-standing fact, not a technical one, and AI has no path to changing it. A model can't hold a license or go to court.
  2. Judgment under novel, unmodeled conditions stays scarce. AI is strong at pattern-matching against everything ever coded into a standard; it's weak exactly where field failures originate — unusual soil behavior, as-built deviations, interactions the code never anticipated. That gap is where senior engineers still earn their fee, and it's widening in relative importance as the "passes the checks" tier gets automated.
  3. Insurers still need to trace discretionary decisions to a party. E&O pricing depends on identifying who could have caught the error. That underwriting logic doesn't change just because a draft came from a model faster — it changes only once "who exercised discretion over what" becomes harder to reconstruct (see NEW, below).
  4. Physical verification in the field is still a human, physical act. Site inspection, load testing, materials verification — none of this became abundant. AI can flag what to check; it can't do the checking.

New axioms

  1. Reconstructing what the AI assumed, after the fact, when a fast and cheap process leaves a thin paper trail. When models were slow to build, the build process itself generated a record of decisions. When generation is instant, firms haven't yet built the habit of logging why the model made the choices it did — so post-failure investigations may have less to examine, not more, even though more models are being produced.
  2. Calibrating how much independent scrutiny a model needs when scrutiny no longer scales with how hard the model was to make. If ten AI-assisted models can be produced in the time one used to take, and each still needs the same verification depth, review capacity becomes the bottleneck at a volume nobody staffed for.
  3. Deciding whether "the engineer directed the AI's assumptions" is a legal fiction or a real fact, at scale. As AI proposes more of the model's substance, the sign-off ritual increasingly certifies judgment the signer didn't fully originate — and nobody has yet had to litigate where that line sits.

Where it breaks

Firms still budget peer review as a slice of design-time cost (INVALID) right as the volume of AI-assisted models needing that same review multiplies (NEW) — review capacity quietly becomes the real constraint on how many projects a firm can safely take on, while everyone's cost model still says design was the expensive part.

Sign-off still legally presumes the stamped engineer authored the model's assumptions (INVALID), but the paper trail proving what was actually reviewed versus AI-proposed is thinner than ever because generation got fast (NEW) — the first field failure with a heavily AI-assisted model will be an accountability case with weak evidence on both sides, decided on a legal fiction nobody's tested yet.

Related axioms

Other axioms