Deflection Is Not Resolution
In February 2024, Klarna announced that its AI assistant had handled 2.3 million conversations in a month — the work of seven hundred full-time agents, resolving issues in under two minutes instead of eleven. It became the most-quoted number in AI customer service, and for a year, every support leader on earth was asked by their CEO why they hadn’t done the same.
Fifteen months later, Klarna’s CEO told Bloomberg the company was hiring human agents again. Prioritizing AI-driven cost cuts, he said, had produced “lower quality” — and customers should always be able to reach a person.
I run support operations for a living, and I want to be careful here, because the lazy reading — “AI support failed” — is as wrong as the hype was. AI handles an enormous share of support volume well, and the economics are real. The lesson is narrower and more useful: the metric that made the headline was deflection, and deflection is not resolution.
Deflection — or containment, the polite word — measures what never reached a human. That is all it measures. A customer who asked the bot, got a confident wrong answer, gave up, and quietly decided to churn is a successful deflection. A customer who rephrased four times and then found the support email is a successful deflection. The dashboard reads 75% contained. The truth could be anywhere.
What most teams get wrong follows directly. They celebrate the containment percentage because it’s the number the vendor’s dashboard leads with — and the vendor chose that number for a reason. They let the tool define “resolved,” usually as “conversation ended without escalation,” which bakes the failure mode into the definition. And they point their QA program at escalated conversations — the ones a human already saw — while the deflected pile, where the silent failures live, goes unread. This is Goodhart’s law running at machine speed: the measure became the target, and the target is now being hit by a system that never gets tired.
Here is the operating regime that actually tells the truth:
Track recontact, not containment. The headline number should be the seven-day recontact rate: how many “handled” customers came back about the same issue, on any channel. Resolution that doesn’t survive a week wasn’t resolution. Use matched cohorts — multi-issue customers make raw rates noisy.
Survey the deflected cohort separately. CSAT averaged across all conversations lets good human saves mask bad bot answers. The question is the satisfaction of customers only the AI touched.
Read the transcripts. A monthly audit — at least a hundred conversations, stratified across deflected, escalated, and recontacted — read by humans with the authority to change things. Audits like this have a way of finding failure modes the dashboard has no category for. (Read ten tickets a week is this lesson’s founder-sized sibling.)
Define resolution in the contract. If you’re a vendor, your client defines “resolved,” not your tool. If you’re the client, write it down: issue closed, verified by recontact silence and customer confirmation. Whoever controls the definition controls the number.
The research is on the side of honesty here. The best field study we have found that AI assistance genuinely raised resolution rates — most for novice agents. The technology resolves things. That is exactly why measuring honestly is safe: the true number is good enough that you don’t need the flattering one. The companies that get hurt are the ones that let the flattering number compound silently — paying down the gap later in churn, brand, and a public reversal.
The deeper principle, which is the thesis of this whole site applied to one queue: the model’s answer is the cheap part. Deciding what counts as an answer — that’s the part you can’t buy, and the part your customers are actually paying for.
Sources & further reading
- Klarna’s AI assistant handles two-thirds of chats in month one — the original claim, from the source.
- Klarna turns from AI to real-person customer service — the other half of the story.
- Generative AI at Work — what AI actually does to resolution, measured properly.
- ‘Improving Ratings’ — Strathern’s formulation of Goodhart’s law.
- In the archive: Klarna’s claim, Agentforce and the agent turn.