No. 194 / 339

Is manual customer service for order issues (returns, delays) still needed when AI can resolve most fulfillment questions instantly?

The shift

Looking up an order, matching it against return/refund/delivery policy, and generating a correct resolution goes from scarce agent time to abundant and instant — for the routine, well-specified query ("where's my order," "how do I return this," "why is my refund late"). The flip is real for the case that has a clean answer in the data; it does not reach the genuinely-wrong order, the discretionary cost decision, or the physical-world exception.

The axioms

  1. Resolving a fulfillment query requires a human to look up the order, interpret the policy, and answer — scarce agent time per contact.
  2. Handling an upset customer at a real service failure requires a human to de-escalate and reassure — scarce emotional labor and standing.
  3. Issuing a refund, replacement, or goodwill credit is a discretionary decision that costs money and needs an owner — scarce authority to spend against margin.
  4. Resolving a physical-world exception (lost, damaged, stolen, suspected fraud) requires a human to investigate and coordinate off-system parties — carrier, warehouse, payment processor — scarce ability to act outside the order record.
  5. An accountable human owns the final resolution the customer relies on — scarce answerability when the call is wrong.
  6. Frontline agents build judgment by handling a volume of easy cases before they meet the hard ones — scarce, gradual on-the-job training.

Invalid axioms

  1. Resolving a fulfillment query requires a human to look up the order and interpret policy. Order status, tracking, return eligibility, refund timing — these are lookups against structured data plus a stated policy, exactly what a tool-using model does well and instantly, at any hour and any volume. Habit-trap: support is still staffed and forecast by ticket volume, so teams plan headcount for query growth that should now be flat-to-falling, and measure the org by contacts handled rather than by the residual of cases automation couldn't close.

Unchanged axioms

  1. Handling an upset customer at a real service failure needs a human. When a wedding gift didn't arrive or a high-value order is lost, the customer isn't asking for information — they want someone with standing to acknowledge the failure and commit to fixing it. A model can produce the words; it can't hold the relationship or carry the weight of the promise. The exception, not the bulk, is where this bites — and it's the exception that decides whether the customer stays.
  2. Issuing a refund or goodwill credit is a discretionary, cost-bearing decision that needs an owner. Deciding to eat the cost of a replacement, waive a restocking fee, or comp a delay is a judgment against margin and precedent, with no ground truth to pattern-match. AI can propose the credit; someone still owns spending real money and the precedent it sets. (Fast-moving: retailers are already handing bounded refund authority to agents for low-value cases. The scarce part isn't the routine sub-threshold refund — it's the novel, high-value, or precedent-setting call.)
  3. Physical-world exceptions can't be resolved from the order record. A parcel marked delivered but missing, a damaged item, a "did the customer or the courier lie" fraud case — resolving these means investigating and coordinating parties outside the system: opening a carrier claim, checking warehouse footage, holding a payment. The model can draft the claim and route it; it can't make the carrier or the warehouse act, and it can't establish what physically happened.
  4. Someone accountable owns the resolution the customer relies on. A model isn't answerable for a wrong refusal or a mistaken refund; the merchant still owns the chargeback, the regulator's letter, and the lost customer. Liability didn't get cheaper because the answer got faster.

New axioms

  1. Frontline agents now inherit only the hard, emotional cases with no easy warm-up. The routine tickets that used to train new agents into judgment are gone to automation, so the human queue is all exceptions, escalations, and angry customers. We must solve for how agents build the judgment those cases demand when the on-ramp of simple tickets that used to build it no longer exists — and for the burnout of a job that is now nothing but the hard 10%.
  2. A confident-wrong AI refusal can lose a customer with no human in the loop. An agent that plausibly but incorrectly denies a valid return, misreads a delay policy, or refuses a legitimate refund fails silently — the customer just leaves, and the failure looks like a resolved ticket. We must solve for who catches the wrong "no," since the old backstop (a human read every contact) is gone and the metric (contact closed) actively hides it.
  3. Mis-resolution at volume is a new verification problem. When one agent handles millions of contacts, a systematic error — a policy edge case read wrong, a refund rule applied too generously or too tightly — replicates across every affected customer before anyone notices. We must solve for sampling and auditing AI resolutions at a scale where reading every one is impossible and the errors are individually plausible.
  4. The physical exception is now the whole remaining job, but it's routed as if it were a fraction of it. With routine queries automated away, the residual human workload is disproportionately the messy, off-system cases — yet staffing, tooling, and comp are still built around a mixed queue of easy-and-hard. We must solve for a support function whose entire remaining volume is the part AI can't touch.

Where it breaks

Support is still measured by contacts handled and forecast by ticket volume (invalid axiom: each query needs a human), while the metric that used to double as quality control — a human saw every contact — is exactly what automation removed, so a confident-wrong AI refusal closes as a successful ticket and the lost customer never shows up in the numbers (new problem: who catches the wrong "no"). The org optimizing for deflection rate is rewarding the failure mode it can't see.

A second collision: teams still assume agents grow judgment by working up from easy cases (invalid axiom: gradual on-the-job training on routine tickets), but automation took the entire easy tier, leaving a human queue that is all exceptions and escalations (new problem: only the hard, emotional cases remain). The people expected to own the goodwill decision and the lost-high-value-order call are the same people who no longer have any low-stakes reps to build that judgment on.

Related axioms

Other axioms