No. 223 / 339
What shifts for the HVAC technician when load calcs and error-code reads are effectively free?
The shift
Two scarce cognitive goods collapse at once: Manual J-style load sizing and error-code interpretation. Both were gated by a trained tech's time and reference knowledge; AI sizing tools now produce a load calc from house inputs in seconds, and a homeowner can photograph a fault code and get a plausible diagnosis before the truck arrives. What doesn't move is anything requiring a body in the mechanical room — commissioning, airflow, the physical install and repair, and a name on the job when the system underperforms.
The axioms
- Sizing a system correctly requires a trained tech who can run the load calc — scarce reference knowledge and hours.
- Reading a fault code and mapping it to a cause requires the tech's diagnostic knowledge — scarce interpretation, often the reason for the trip.
- The customer can't self-diagnose, so the diagnostic visit is billable and the knowledge asymmetry justifies the trade rate.
- A system that runs is a system installed and commissioned right — airflow, static pressure, refrigerant charge, duct reality — scarce physical work that no calc substitutes for.
- Whatever the calc says, the actual building governs: leaky ducts, undersized returns, real infiltration, room-by-room reality — scarce on-site verification against ground truth.
- Someone is accountable when the system doesn't heat, cool, or last — a name on the install and the warranty callback.
- The customer trusts the tech's judgment about what to buy and whether the fix held — scarce standing built over service calls.
Invalid axioms
- Sizing correctly requires a trained tech who can run the load calc. Manual J is a deterministic procedure over building inputs — exactly what software has done for years and what AI tools now do faster, more consistently, and without the rule-of-thumb oversizing techs default to under time pressure. The calc itself is no longer the scarce good. Habit-trap: shops still treat "we do a proper load calc" as a differentiator and price it as expert labor, when the sizing number is now close to free and arguably more consistent than a rushed tech's tonnage-per-square-foot guess.
- Reading a fault code and mapping it to a likely cause is scarce diagnostic knowledge worth a trip. Code-to-cause lookup is pattern-matching against everything ever documented — the LLM's home turf. Customers photograph the board, describe the symptom, and get a ranked list of causes before calling. Habit-trap: the "diagnostic fee" is still often priced as if interpreting the code were the value, when interpretation has largely leaked to the customer's phone; what remains billable is confirming it on the equipment, not naming it.
- The knowledge asymmetry justifies the trade rate. The rate rested partly on the customer not knowing what the code meant or what size unit they needed. That specific asymmetry is thinning fast. Habit-trap: pitching the value as "we know things you don't" invites a customer who arrives already holding a plausible (and possibly wrong) AI answer and treats the tech as a second opinion rather than the authority.
Unchanged axioms
- A system that runs is a system installed and commissioned right. The load number tells you the equipment; it says nothing about whether the charge is correct, static pressure is in range, airflow per ton is delivered, the condensate drains, and the ducts actually move the air the calc assumed. This is measurement and mechanical work at the equipment — refrigerant gauges, manometer, hands on sheet metal — and it stays scarce and physical regardless of how good the calc is. Most systems that "don't work right" were sized fine and commissioned poorly.
- The actual building governs, and only on-site measurement reveals it. An AI sizing tool takes the inputs it's given and returns a confident number; it can't see the return that's half-blocked, the duct run crushed in the crawlspace, or the addition the homeowner didn't mention. Verifying the calc against the real house — and catching where the inputs were wrong — is ground-truth work the model can't do. This is where "confidently wrong AI sizing" gets caught, if anyone's checking.
- Someone is accountable when the system underperforms. The AI that sized the unit isn't liable when the second floor never cools; the installer's name is on it. Warranty, callbacks, code compliance, and the standing to be the party who answers for it stay human and stay on the trade. A model can produce the number; it can't own the outcome.
- Trust is built at the equipment, not at the calc. The customer's confidence that the fix held and the install was done right comes from the tech being physically competent and answerable — reinforced, not replaced, when the customer can now sanity-check claims against AI. Trust shifts from "trust my knowledge" to "trust my hands and my accountability," but it stays scarce.
New axioms
- Customers arrive anchored on an AI sizing or diagnosis that ignored the real house. The homeowner has a number and a named cause before the truck arrives, delivered with more confidence than uncertainty. When on-site reality contradicts it — the calc that assumed tight ducts, the code that had a mechanical cause the phone couldn't see — the tech now has to un-sell a confident-wrong answer, which is harder than informing someone who knew nothing. The scarce act moves from producing the answer to correcting a plausible wrong one the customer already believes.
- The correction work is real labor and nobody's paying for it yet. Verifying the AI's calc against the building, catching the bad input, and re-diagnosing past the customer's phone-answer is skilled effort — but it looks like "arguing" to a customer who thinks the answer is already settled and free. Shops that unbundled and stopped charging for "the calc" or "the diagnosis" haven't priced the verification-and-correction work that replaced it.
- The blind-spot gap is invisible until the system underperforms. A confidently wrong AI size (right on paper, wrong for the duct system) or a missed mechanical cause behind a fault code produces a job that passes on inputs and fails in the house — and the failure surfaces months later as a callback, by which point it reads as the installer's fault, not the calc's. The residual gap between the confident number and the real building is now the trade's liability, uncaught unless a tech is explicitly checking for it.
Where it breaks
Shops that unbundle and stop charging for "the load calc" and "the diagnosis" (INVALID: those are near-free now) haven't repriced the verification-against-the-real-building work that replaced them (NEW: correcting confident-wrong AI is unpaid labor). The tech now does the harder job — un-selling a customer's confident wrong answer and catching where the inputs lied — for a fee structure built to charge for the answer, not the correction. And when a right-on-paper size fails in a real duct system, the AI that produced the number carries no liability; the installer's name does (STILL HOLDS: accountability stayed on the trade), so the party least responsible for the bad input owns the callback.
Fast-moving flag: how far diagnosis genuinely leaks to the customer depends on multimodal + tool-use trajectory — a phone reading a board, ingesting the manual, and walking a homeowner through basic checks is close now and improving. Physical commissioning is not on that curve; the gap between "AI can name the cause" and "AI can charge the system" is the whole audit, and it widens rather than closes as the diagnostic side gets better.
Related axioms
Other axioms
Engineering
What changes for ML engineering with AI?
Cybersecurity
Who's liable for a breach an AI security agent missed or misclassified?
Finance
What's a financial advisor for once portfolio construction and tax-loss harvesting are commoditized?
Engineering
What changes for DevOps with AI?
Society
Does AI intake and documentation give social workers back time for care, or just raise the caseload expectation?
Engineering
What gets measured as engineering productivity now that lines-of-code and PR count are gamed by agents?