No. 314 / 339
Should "average handle time" still be the metric we optimize for when AI handles the easy tickets and humans only get the hard ones?
The shift
Handling the fast, high-volume, low-variance tickets goes from scarce (agent minutes rationed against a queue) to abundant — an AI clears them instantly and near-free. What's left in the human queue is the residue: the hard, novel, high-emotion, high-stakes cases AI couldn't close. AHT was built to ration scarce agent minutes across a queue dominated by easy tickets; that queue no longer exists on the human side.
The axioms
- Support cost is roughly agent-minutes times ticket volume, so minutes-per-ticket is the lever that moves the budget (scarce agent time).
- Faster handling means more tickets closed per agent per hour, which means lower cost per contact (scarce throughput).
- AHT is a decent proxy for productivity because tickets are similar enough that a faster agent is a better agent (abundant near-uniform tickets — the average is meaningful).
- You can't manage what you can't measure, and handle time is cheap and objective to measure, so it becomes the target (scarce, cheap, objective measurement).
- The queue is the enemy; keeping wait times and backlog down is the operational job (scarce agent time against abundant demand).
- What you actually care about — did the customer's problem get solved, do they still trust us — is expensive and slow to measure, so a fast throughput proxy stands in for it (scarce measurement of quality/trust).
- The human is fungible against the queue: any agent can take any ticket, so the metric is per-agent averaged (abundant interchangeable labor on interchangeable tickets).
Invalid axioms
- AHT is a decent proxy for productivity because tickets are similar enough that faster equals better. This rested on a queue full of near-uniform tickets where the average was meaningful. AI has drained exactly the uniform tickets, so the human queue is now high-variance by construction — a 4-minute case and a 90-minute case aren't points on one distribution. The average of a bimodal-gone-unimodal-hard queue measures almost nothing. Habit-trap: dashboards still show a single AHT number and still trend it, comparing this quarter's all-hard queue against last quarter's mostly-easy one and reading the inevitable rise as agents getting slower.
- Faster handling means lower cost per contact, so minutes-per-ticket is the budget lever. Speed on the easy tickets was where the cost was, because that's where the volume was. AI took the volume, so squeezing human minutes now moves a small and shrinking share of total cost. Habit-trap: WFM and ops teams still set AHT targets and coach to them as the primary cost lever, optimizing the residual while the real cost equation has moved to AI accuracy, escalation rate, and containment.
- The human is fungible against the queue; measure per-agent averaged. When any ticket could go to any agent, averaging made sense. The hard queue rewards specialists, deep context, and staying with one case — the opposite of interchangeable. Habit-trap: routing and scoring still treat agents as uniform capacity and rank them on a shared average, penalizing the person who takes the gnarly cases nobody else can close.
Unchanged axioms
- You measure so you can manage, and handle time is still cheap and objective to capture. The instrumentation isn't wrong; the interpretation is. Time-in-case is still worth logging — as a diagnostic (which case types are blowing up, where agents are stuck, where tooling is slow), not as a target. It just stops being the thing you optimize. The scarce, expensive thing to measure — did this actually get resolved, does the customer still trust us — is still scarce and still not captured by a clock.
- Someone accountable has to own whether the hard case was actually resolved. AHT never measured this and still can't. On the residue queue, resolution quality, correctness under ambiguity, and whether the customer walks away trusting the company are the whole job — and those rest on human judgment and accountability, which AI hasn't made abundant. This is the metric that should move to the center; the fact that it's hard to measure is exactly why AHT was allowed to stand in for it.
- The customer relationship on a hard case is carried by a person, not a fluent answer. A customer who has already been through the AI and still has a problem is, by definition, someone the fast path failed. What resolves them is judgment, ownership, and the sense that a person is genuinely on it — none of which a stopwatch rewards, and some of which a stopwatch actively punishes.
New axioms
- A metric that stops punishing humans for handling only-hard tickets. Every legacy throughput metric — AHT, tickets-per-hour, contacts-per-agent — now reads worse precisely because the work got harder, not because performance dropped. Left unchanged, they systematically penalize the humans doing the highest-value work and reward gaming (fast punts, premature closes, bouncing hard cases back to the queue). The org needs a measure of the hard queue that survives the easy tickets being gone.
- Measuring resolution quality and trust directly, now that speed is free and no longer the constraint. AHT was tolerated because quality was expensive to measure and a throughput proxy was cheap. With speed effectively solved on the easy path, quality-of-resolution and retained trust are the only things left worth optimizing — but most orgs have no live instrument for them beyond a lagging CSAT. Building a fast, trustworthy read on "did this actually get solved well" is the open problem. (Whether AI-graded resolution quality can fill this gap is moving fast — plausible within a year or two, and worth flagging as a call that hinges on it.)
- Goodharting a throughput proxy when throughput is no longer the human's job. The moment you keep AHT as a target on a hard-only queue, you've picked a proxy whose only remaining effect is distortion — it can't make the cheap tickets cheaper (AI already did) and it actively degrades the hard ones. The problem is organizational muscle memory: the proxy outlived the scarcity it was proxying for, and killing a familiar number is harder than adding a new one.
- What "productive" even means for a support human once volume isn't their contribution. If a human closes six hard cases a day and that's the whole valuable output, capacity planning, staffing ratios, and performance reviews all lose their denominator. The org has to define output in terms of hard-case resolution, escalation prevention, and AI-oversight quality — categories most support orgs don't measure at all yet.
Where it breaks
The dashboard still trends a single AHT number as the health metric (INVALID #1) at the same moment the only tickets reaching humans are the hard ones (NEW #1) — so the metric mechanically rises every quarter as AI containment improves, and an org reading it literally will conclude its remaining agents are getting slower and cheaper to cut, when what's actually happening is they've been handed nothing but the hard cases. The better the AI gets, the worse the humans look on the old number.
Speed was kept as the target because quality was too expensive to measure (INVALID #2 relying on the old proxy) right when quality-of-resolution becomes the entire human job (STILL HOLDS #2, NEW #2) — the org is still optimizing the one thing that's now free and still not measuring the one thing that's now everything.
Related axioms
Marketing
What changes for marketing and advertising with AI?
Marketing
What changes for sales with AI?
Marketing
What changes for customer success with AI?
Marketing
What changes for customer support with AI?
Marketing
Do we still need product marketing to translate features into positioning if AI drafts messaging from changelogs?
Marketing
What's an ad agency's fee structure for once media buying and creative optimization run on autopilot?
Other axioms
Legal
Is the paralegal role dead, or does it just move upstream into AI-output verification?
Engineering
Who's accountable for data quality when AI both generates and validates its own training data?
HR
Does compensation benchmarking still need a dedicated analyst when AI can model market pay in real time?
Society
What changes for clergy and religious institutions with AI?
Industries
What changes for mining and extractive industries with AI?
Engineering
What changes for work when AI safety, alignment, and red-teaming become core functions?