Shep Hyken has highlighted a problem for companies introducing AI into customer service: the figures used to declare success may say little about whether customers are getting a better service
In a Forbes article published in late June 2026, Hyken examined research showing that many contact-centre leaders regarded their AI projects as successful even when projects were delayed, over budget or causing customer friction.
The numbers create an awkward question. If a system saves money or reduces the number of people reaching an agent, does that mean the customer’s problem was solved, or only that it disappeared from the company’s view?
Success, despite the warning signs
The research came from Laivly, a company that supplies AI technology to contact centres. It surveyed 200 leaders responsible for recent customer-service AI projects involving external vendors.
Sixty-five per cent called their most recent AI initiative successful. Yet 43% of projects were delayed or stalled, while 53% had exceeded their budgets. (Laivly)
Another 28% said they had lost revenue because AI could not deal with customer complexity. A further 20% believed revenue was being lost but said they could not quantify the damage. (Laivly)
What is the metric measuring?
Companies need internal measures. They need to know whether a project launched, how much it cost, how often it is used and whether it reduces pressure on staff.
But those measures can become misleading when they are treated as proof that the service improved.
A chatbot may handle thousands of enquiries without showing whether the answers were useful. A call may be deflected from a human agent without revealing whether the customer found a solution or simply gave up.
When deflection looks like success
One common measure is containment: the proportion of customer contacts completed without reaching a person.
That can be useful when automation genuinely answers a simple question. Most of us don’t need a human conversation to check an opening time or find a standard form.
The problem is that containment may also include people who abandon the route, accept an incomplete answer, search elsewhere or start again through another channel.
The work may not disappear
When automation fails, the work often moves from the company to the customer.
We try different wording, reopen the chat, search through help pages and repeat information that has already been supplied. Eventually, we may look for a phone number that is no longer easy to find.
From the company’s side, the original contact may look closed. From our side, the problem is still there, and we are now doing more of the work needed to solve it.
Customer friction has a cost
Laivly reported that 49% of the companies surveyed had experienced increased customer friction linked directly to their AI tools.
Among those companies, 57% said that friction was costing them between 5% and 10% of sales. The figures are self-reported, but they suggest that poor service design may eventually show up in financial results. (Laivly)
This is where the title’s “cost-cutting bonanza†needs some care. The research doesn’t prove that every AI project produces large savings, or that every company is cutting service deliberately.
Cost-cutting may be part of the picture
Laivly found that 78% of companies expected AI savings through reductions in agent numbers, while 44% planned cuts within 12 months.
It also reported that the companies cutting most aggressively were more likely to report customer friction, revenue leakage and higher project costs. (Laivly)
That does not establish that staff reductions caused every problem. It does raise the question of what happens when savings are counted immediately while the effects on access, trust and repeat contact appear later.
Abandonment isn’t resolution
A customer leaving a chatbot does not necessarily mean the chatbot completed its job.
Perhaps the answer was useful. But the customer may also have lost patience, decided the amount at stake was not worth pursuing or concluded that the company had made further contact too difficult.
Unless organisations investigate what happens afterwards, they may mistake disappearance for satisfaction.
Can customers still reach a person?
Hyken’s wider argument has often been that businesses should take an “AI-first, not AI-only†approach, using automation while keeping people available when the technology reaches its limits.
That idea fits closely with the question raised by these figures. A company may say human help remains available, but the useful test is how easily someone can actually reach it.
How many steps does the handover take? Does the customer have to repeat everything, and does the person receiving the case have enough context and authority to get something done?
A reply is not always help
Automated systems can respond almost instantly, which may make an organisation appear highly responsive.
But speed tells us little if the answer is irrelevant, circular or based on a misunderstanding that the customer cannot easily correct.
At ReplyResearch, this is the kind of gap we might describe as contact theatre: enough visible activity to suggest reachability, without a dependable route to anyone able to deal with the problem.
Make room for the innocent explanation
Some weak AI projects may reflect hurried deployment, fragmented software, unrealistic expectations or poor integration rather than a deliberate attempt to block customers.
Laivly found that 56% of companies were using more than three AI tools. Organisations with larger, fragmented technology stacks were more likely to report delays and increased friction. (Laivly)
That may explain part of the problem, but it doesn’t remove the effect on customers. Whatever caused the failure, the extra effort still lands on the person trying to get help.
What should companies count?
Launch dates, contact volumes and savings do not need to disappear. They need to sit beside measures that show what happened to the customer.
Those could include repeat contacts about the same issue, failed handovers, contradictory answers, abandoned journeys and the amount of information people have to provide again.
Companies could also measure whether the eventual answer resolved the original problem, rather than assuming that an automated interaction was successful because it ended.
Test it from the outside
The most revealing tests may begin with problems that do not fit the standard script.
Ask an unusual question. Challenge an incorrect response. Try to reach a person, and see whether the conversation history follows the customer through the handover.
Then test whether the person at the end can actually act. That is closer to the service customers experience than a dashboard showing how many contacts never reached an agent.
The research has limits
Laivly sells contact-centre AI technology, and its report is based on responses from 200 leaders involved in recent projects using external vendors.
The findings describe what those leaders reported. They do not independently measure the experiences of customers using each system. (Laivly)
The figures are therefore better treated as evidence of a tension inside AI deployment than proof that every company is measuring success badly.
What happens after the savings?
AI may reduce repetitive work, help staff find information and resolve straightforward enquiries faster. Used well, it could make human support more effective rather than simply smaller.
The risk is that companies count the immediate savings while failing to notice the customers who are repeating themselves, abandoning complaints or quietly deciding to leave.
Hyken’s warning is useful because it asks what the success metric leaves out. If today’s cost-cutting bonanza depends on customers carrying more effort and receiving less help, how long will it be before those savings return as tomorrow’s lost customers?

Footnote Zone for Wrong AI Metrics: Today’s Cost-Cutting Bonanza Could Be Tomorrow’s Lost Customers
Disclosure: The diagnostic tools referenced below were developed by NokNok, a specialist in online responsiveness tool design.
This Footnote Zone uses NokNok’s diagnostic toolkit to examine whether AI-driven customer-service savings conceal contact barriers, unresolved enquiries, poor automated responses, and broken escalation journeys.
- Email Finder: As organisations direct customers towards automated systems, web forms, and difficult-to-exit support journeys, accessible email addresses and alternative human contact routes may become harder to find. Email Finder scans an organisation’s website and related public-facing materials for published email addresses, then reports missing routes, discrepancies, structural deficiencies, and other contactability gaps.
- Reply Radar: AI containment figures may classify an interaction as completed even when the customer abandons the process, searches for another channel, or joins an understaffed human queue. Reply Radar deploys targeted test emails and quantitatively measures reply rates, response latency, consistency, and related responsiveness benchmarks to determine whether dependable human response operations remain available.
- Compliance Sniffer: Automated systems may produce fast but irrelevant, circular, incomplete, or evasive answers that appear responsive without resolving the customer’s problem. Compliance Sniffer analyzes incoming responses against objective benchmarks for quality, clarity, relevance, escalation, and compliance, helping distinguish meaningful assistance from automated contact theatre.
- Mystery Shopper: Customers may encounter obstructive forms, repeated requests for information, failed handovers, hidden escalation routes, or agents who lack the context and authority to act. Mystery Shopper executes a comprehensive end-to-end responsiveness UX audit, testing how a real user experiences the organisation’s contact, response, handover, and escalation pathways.
Disclosure: The diagnostic tools referenced in this Footnote Zone were developed by NokNok, a specialist in online responsiveness tool design. ReplyResearch may use NokNok tools, resources, or analysis when preparing coverage, while retaining responsibility for its editorial decisions, including what topics to cover, what sources to cite, and how stories are presented. Read the full ReplyResearch Collaborative Disclosure Policy [here].

Sources and relevant reading for Wrong AI Metrics: Today’s Cost-Cutting Bonanza Could Be Tomorrow’s Lost Customers
- “The Most Dangerous AI Metric Is the One That Says You’re Successful†– Shep Hyken, Forbes
28 June 2026
Read the article
This is the commentary directly discussed in the article. Hyken examines the contradiction between organisations declaring AI projects successful and evidence of delays, excess costs, customer friction and lost revenue. It supports the central argument that deployment, containment and cost reduction should not automatically be treated as evidence of improved customer service. - “65% of Contact Center Leaders Call Their AI Successful. Yet 43% of Projects Are Delayed or Stalled†– Laivly
29 June 2026
Read the research announcement
This is the primary source for the survey figures cited in the article. It reports that 65% of customer-experience leaders regarded their latest AI project as successful, although 43% of projects were delayed or stalled, 53% exceeded their budgets and 28% caused revenue losses because the systems could not handle customer complexity. It also provides context for the discussion of staffing reductions, fragmented technology and customer friction. - “Gartner Predicts 50% of Organizations Will Abandon Plans to Reduce Customer Service Workforce Due to AI†– Gartner
10 June 2025
Read the announcement
Gartner predicts that half of organisations planning substantial AI-related customer-service workforce reductions will abandon those plans by 2027. The finding supports the article’s warning that companies may count immediate staffing savings before understanding whether automated systems can reliably manage complex customer needs. It also reinforces the case for retaining accessible human support. - “Agentic AI and Human-in-the-Loop Interventions: Field Experimental Evidence from Alibaba’s Customer Service Operations†– Yiwei Wang and others
14 May 2026
Read the research paper
This field experiment found that AI reduced average chat duration but lowered customer ratings for conversations handled by the system. It also found that the effectiveness of human intervention depended on the nature of the AI failure and how early the intervention occurred. The research relates closely to the article’s argument that speed and shorter interactions do not necessarily indicate resolution or customer satisfaction, and that well-designed escalation to a person is essential. - “Silent Abandonment in Text-Based Contact Centers: Identifying, Quantifying, and Mitigating Its Operational Impacts†– Antonio Castellanos and others
15 January 2025
Read the research paper
This study examines customers who leave text-based support conversations without formally ending them. Across the organisations studied, silent abandonment made it difficult to determine whether customers had received help or simply disappeared. The research directly supports the article’s distinction between containment and resolution and its warning that an ended interaction may conceal abandonment rather than success. - “AI Shakes Up the Call Center Industry, but Some Tasks Are Still Better Left to the Humans†– Associated Press
September 2025
Read the article
This report describes how contact centres are using AI for routine work while continuing to rely on human agents for complicated or sensitive cases. It provides broader industry context for the article’s “AI-first, not AI-only†argument and shows why automation works best when it assists employees or handles straightforward requests rather than becoming a barrier to human help. - “The 5 Fastest Ways to Get Past AI Customer Service Chatbots†– Tom’s Guide
June 2026
Read the article
Based on practical tests of customer-service systems, this article documents the effort sometimes required to escape automated conversations and reach a person. Its findings illustrate the customer-side workload described in the article: repeating requests, using particular trigger words, changing channels and learning how an organisation’s escalation system operates before receiving human assistance. - “AI to Human Escalation: Designing Seamless Customer Service Handoffs†– Pertama Partners
12 December 2025; updated 17 June 2026
Read the article
This practical analysis focuses on escalation triggers, preserving conversation history and ensuring that human agents receive the context needed to continue an interaction. It relates to the article’s questions about whether customers must repeat information, how many steps a handover requires and whether the eventual agent has enough knowledge and authority to resolve the original problem.
