
In short
Chatbot metrics become misleading when they treat silence, abandonment, fast acknowledgments, or closed conversations as successful resolution; stronger measures track repeat contact, verified completion, useful responses, and what the customer does next.
A chatbot dashboard can be green while customers are still stuck.
The problem is not always bad data. Sometimes the metric itself rewards the wrong event: a customer goes quiet, an automatic acknowledgment stops the clock, or a conversation closes without proof that anything was fixed.
We asked one question: Which customer service metric did you stop trusting, and what replaced it?
Twelve support and business leaders answered. Across eight familiar metrics, the replacements kept pointing away from the chat window and toward what the customer did next.
1. Chat resolution rate: silence counted as solved
“A woman who walks away mid-wax with wax cooling on her leg looks identical to a woman whose question got answered.”
Dan McElwee | Head of Retail, Tress Wellness
Read full expert response
Chat resolution rate was the number I stopped trusting. Our automated flow counted a conversation as resolved the moment the customer stopped replying, and a woman who walks away mid-wax with wax cooling on her leg looks identical to a woman whose question got answered.
What I read instead is repeat contact tied to the same order. If someone comes back within a week on the same purchase, that first chat did not resolve anything, whatever the dashboard said. I pair that with the return rate on orders that touched support, and with what those same buyers write in reviews about irritation or skin reaction.
Instrumenting it was unglamorous. I had chats tagged by topic and matched to the order ID, then pulled the repeat-contact and return columns weekly. No new platform, no analyst on payroll. The real cost was a few hours a week of my own time reading the threads, which is the part nobody wants to price.
The reading changed the flow. Questions about skin reactions and burning no longer get an automated answer at all, they route to a person. And the pre-wax and post-wax steps that customers kept re-asking about now appear in the first reply, before they ask twice.
“Silence isn't satisfaction.”
Rory Keel | Owner, Equipoise Coffee
Read full expert response
The metric I fired was chat resolution rate. Our automated chat kept reporting sky-high resolution numbers, and it looked beautiful right up until I noticed the pattern: someone asks whether our Ethiopian Yirgacheffe leans bright or smooth, the bot half-answers, the customer goes quiet, and the system logs a win. Silence isn't satisfaction. In roasting terms, that's like calling a batch done because the beans stopped crackling.
What replaced it: repeat-contact rate within seven days, tied to actual orders. If a customer asks about brewing our Mexican La Laja Honey and never messages again, that could mean delighted or defeated. But if they come back with the same question, or ask a human to restart the conversation, the bot failed. And if they reorder the Cavaliers Blend, that's a signal I can take to the bank.
How I instrument it: every conversation gets tagged by topic, then I cross-reference those tags against repeat contacts and order history. It's deliberately low-tech. We're a small-batch roastery, so I didn't buy enterprise tooling; the real cost was an hour a week of honest reading plus a simple spreadsheet. What it bought me was the truth: chat handles shipping and order questions well because those have clean answers, but it struggled on brewing science, which is exactly why we invest in our educational blog instead of pretending a bot can taste coffee.
My advice for your readers: measure what a customer does after the conversation, not what the software claims during it. Deflection counts bodies leaving the room. Repeat contact and reorders tell you whether they left happy. That's the same balance we chase in every roast profile at Equipoise Coffee, and it applies to dashboards too. If a metric can't tell the difference between a solved customer and a gone one, it isn't a metric, it's a mirror.
Both replacements ask the same question after the chat closes: did the customer return with the same problem? That requires linking conversations to a topic, order, or account and accepting that a person still has to read the ambiguous cases.
2. Deflection rate: leaving is not the same as succeeding
“A customer who exits a chat out of frustration is regarded as deflected, but usually not resolved.”
Pratik Singh Raguwanshi | Manager, Digital Experience, CXrove
Read full expert response
The deflection rate is an extremely dangerous vanity metric in automated customer support because it actually encourages the chatbot to be unintuitive rather than helpful to the users. With fifteen years of experience in BPO operations, I have seen operational teams cheering at the same time, because they had managed to achieve a deflection rate of 70%, while scores of customers were turning to voice support agents after being unable to get help from the chatbot. We came to the conclusion that a customer who exits a chat out of frustration is regarded as deflected, but usually not resolved, and they are just waiting for the lunch break to call the help desk at some later point and complain there.
We stopped trusting the deflection rate as the only metric and started thinking about using the Resolution Efficiency metric instead, which indicates whether the task was successfully completed without any follow-up actions through any available channel for the following 24 hours. We had to implement a lot of changes in our approach to data collection in order to make use of this new metric. One of the changes was abandoning fragmented reports on chat sessions in favor of integrating chat information with CRM and telephony data.
Using the new metric did not require a lot of money to be spent on a new software solution, but it did require our analysts to invest many hours of their time. We needed to spend 20 hours in the first quarter just to process the available data and create one unique model, which would connect customers with unresolved issues through chat with other interactions through email or phone.
After we implemented this new approach, our attitude to building chatbots has changed significantly. Instead of using chatbots as obstacles to prevent people from communicating with us, we started thinking about how to optimize our processes in terms of getting the answer as fast as possible. We realized that having a lower deflection rate in light of having no follow-up calls is way more important than reaching a higher deflection rate that leads to a rise in expensive call centre inquiries.
“If your success metric can improve while your customers get angrier, it's not a success metric.”
Runbo Li | CEO, Magic Hour AI
Read full expert response
We stopped trusting deflection rate. It was the most dangerous number on our dashboard because it rewarded silence. A user who rage-quit the chat after two unhelpful bot responses counted the same as someone who genuinely got their answer. For months, our deflection rate looked fantastic. Meanwhile, churn from users who never got real help was quietly eating into retention. We were celebrating a metric that was literally measuring how many people we successfully ignored.
What replaced it is something I call “resolution-to-return.” It measures whether a user who interacts with our automated support comes back and takes a meaningful product action within 48 hours. Not “did they stop asking questions,” but “did they go make a video afterward.” That's the only proof that the support interaction actually worked.
Instrumenting it was straightforward but required stitching together two data sources that most teams keep separate: support session logs and product event streams. We pipe our chat logs into our analytics layer and join them on user ID with a 48-hour attribution window to the next completed video render. The whole pipeline runs on tools we already had. No new vendor, no new headcount. It took me about a weekend to build the first version and maybe four hours a month to maintain.
The moment we switched, our “real” resolution rate dropped from what looked like 78% to around 41%. That was a gut punch, but it was the truth. It told us exactly which bot flows were dead ends. We rebuilt three of our highest-volume automated responses, added clearer escalation paths, and within six weeks that resolution-to-return number climbed to 63%. More importantly, 30-day retention for users who hit support improved noticeably.
The lesson for any CX leader: if your success metric can improve while your customers get angrier, it's not a success metric. It's a vanity metric with a job title. Tie your support measurement to the thing your product exists to do. For us that's rendering a video. For you it might be completing a purchase or launching a campaign. The point is, support didn't succeed until the customer succeeded. Everything else is just counting silence.
A lower deflection rate can be the healthier outcome if fewer customers need to return through phone or email. The useful denominator is not conversations that ended. It is customers who completed the job they came to do.
3. Containment rate: one number hid four different endings
“Containment rate counted someone giving up on the bot as a win.”
Lilach Bullock | AI Implementation Consultant and Fractional CMO
Read full expert response
The one I stopped trusting was containment rate, because it counted someone giving up on the bot as a win. I replaced it with tagging every automated conversation as resolved, escalated, abandoned mid-flow or off-topic, done inside HubSpot's own conversation log rather than a separate dashboard. Going through those tags by hand took the best part of a morning each week, which is the bit nobody budgets for. That review showed billing queries were the ones people abandoned most, so we pulled the bot off billing entirely and routed those straight to a person.
Resolved, escalated, abandoned, and off-topic conversations should not collapse into one containment percentage. The categories expose which work belongs with the bot and which should go directly to a person.
4. First response time: the acknowledgment stopped the clock
“A fast first reply and a resolved problem are not the same thing.”
Rick Elmore | CEO, Simply Noted
Read full expert response
I stopped trusting first response time. We were hitting it easily, our dashboard was green all month, and churn was still creeping up in our small business accounts.
The problem was that a fast first reply and a resolved problem are not the same thing. Our automated chat was firing an instant acknowledgment, which stopped the clock, and then a real person picked it up two hours later. On paper we looked fantastic. In practice the customer sat there.
What replaced it was time to first USEFUL response, meaning the first message that either answered the question or asked something the customer had to answer. We instrument it by tagging outbound messages in our helpdesk as either acknowledgment or substantive, and the timer only stops on substantive. Took our ops lead about a day and a half to set up the tagging rules and roughly two weeks of tuning to stop miscounting canned replies as substantive. No new tooling spend.
The number got uglier immediately. It went from under a minute to about 47 minutes, and that was the honest picture. We restaffed our afternoon coverage and got it to about 12. Support driven cancellations dropped noticeably the quarter after.
Time to first useful response is harder to flatter. It starts the same clock but only stops when the customer gets an answer or a question they can act on.
5. The thumbs up at the end: failed users often never vote
“Someone the automation fails does not rate it. They close the tab and try something else.”
Anton Strasburg | Media Manager, FreeConference.com
Read full expert response
The one I stopped trusting is the rating at the end of an automated chat, the thumbs up or thumbs down. On a free product with a big self-serve base it reads like a health score and it is not one.
The trouble is who answers. Someone the automation fails does not rate it. They close the tab and try something else, and you never hear from them. The people who do rate are the delighted ones and the ones desperate enough to keep hammering the widget until something works. Neither group is the ordinary user. Published benchmarks put post-chat survey response in the low double digits at best, and automated conversations sit at the bottom of that range. Take that as an industry range rather than a figure off my own dashboard.
What I count instead is repeat contact on the same topic inside seven days. Every conversation gets one tag from a short list, keyed to the session or the account, so a second visit lines up with the first. No new tooling. The cost is the tagging discipline, and somebody reading the repeats each month instead of glancing at a chart.
It changes where the effort goes. Rather than a desperate push to lift a score, the work lands on the handful of help pages the automation keeps failing on, and a repeat that stops repeating is something you can check.
A rating is feedback from the people who stayed long enough to give it. Repeat contact adds the missing behavior from those who closed the tab.
6. Reported resolution gave way to verified resolution
“A case is only marked resolved when a post-chat CSAT confirms closure or when no related issue reopens within 72 hours.”
THERY Jean Christophe | CEO, MusaArtGallery
Read full expert response
I stopped trusting chat deflection rate because abandoned or unanswered conversations were counted as resolved, creating a false sense of success. We replaced it with a verified resolution rate: a case is only marked resolved when a post-chat CSAT confirms closure or when no related order/issue reopens within 72 hours.
We instrumented this by linking chat transcripts to order IDs, sending an automated CSAT email after the chat, and flagging any reopened tickets or fulfillment exceptions as unresolved. Implementation used our Shopify/helpdesk stack plus AI-assisted automation to tag and reconcile events, and required configuration changes and a few weeks of analyst time to map events and reports.
The result: fewer false positives and clearer prioritization of automation improvements.
“A closed chat does not equal a satisfied customer.”
Silvia Lupone | Owner, Stingray Villa
Read full expert response
The first time I lost faith in this metric for evaluating customer support was when I realized how quickly I could “fix” a customer's issue with a simple chatbot. On the surface, chat deflection seemed like a brilliant way to measure success. A customer has a question; the automated chat provides an answer; no human is engaged. Success!
However, there is a huge difference between a conversation closing (which is all you know if you're only looking at your metrics) and a customer problem being solved. For example, someone may ask where to go to pick up their dive equipment, about check-in times, or whether we are a good fit for their needs. A chat bot gives the customer an answer and then closes the conversation. However, you have no idea if the customer trusted that answer enough to complete a booking.
What I am interested in now, instead of whether the automated chat answered their questions, is what happened after the conversation ended.
Were customers still browsing through my website? Did customers reach out to us again? Were bookings completed? Or were guests able to leave without ever taking another step?
This seems very basic, however it changed how I view automated chat. I do not want systems that make my numbers look better. I want systems that help a real person decide whether or not to take action.
A closed chat does not equal a satisfied customer. Satisfaction comes from what the customer does after the chat.
Verification can be explicit, through a satisfaction confirmation, or behavioral, through no related reopen and a completed booking or purchase. Either is stronger than treating a closed window as proof.
7. First contact resolution: the older version of the same mistake
“Customers don't care if you responded fast. They care if they have to contact you again.”
Joe Spisak | CEO, Fulfill.com
Read full expert response
We killed “First Contact Resolution” at my fulfillment company around 2019 because it was rewarding our team for closing tickets fast, not solving problems. Our FCR hovered at 82%, which looked great in board decks. Then I started getting calls from brand founders saying “your support is fast but I keep having the same issue.” That's when I realized we were measuring speed to close, not actual resolution.
The breaking point was a DTC apparel brand that opened 14 tickets in 30 days about the same inventory discrepancy. Each ticket got marked “resolved” because our team responded within SLA and the customer stopped replying. FCR said we were crushing it. Reality said we never fixed the root problem with how their ASN data was flowing into our WMS.
I replaced FCR with what we called “Issue Recurrence Rate” - tracking whether the same customer contacted us about the same core problem within 30 days. Sounds simple but it required our support platform (we used Zendesk) to tag issues by root cause, not just topic. We added a custom field for “problem type” with about 40 options, then built a dashboard that flagged any customer who opened multiple tickets with matching tags. Cost us maybe 15 hours of analyst time to set up the Looker dashboard and another 20 hours training the team on consistent tagging.
Recurrence rate sat at 31% when we first measured it. Brutal. But now we had real data. We started routing repeat issues directly to operations managers instead of closing them as “resolved.” Within six months recurrence dropped to 11% and our NPS jumped 18 points. Brands stopped churning over support frustrations.
The insight that changed everything: customers don't care if you responded fast. They care if they have to contact you again. Measuring closure speed just optimizes for making problems disappear from your queue, not from your customer's life.
This example came from human support, not a chatbot, but the failure mode is identical. If closing the ticket is rewarded, the system optimizes the queue. Recurrence asks whether the underlying problem disappeared from the customer’s life.
8. Vendor dashboards and reply counts: activity without proof
“We'd rather check the actual work than trust a dashboard.”
Marcos De Andrade | Founder and Owner, Green Planet Cleaning Services
Read full expert response
The metric I stopped trusting was the resolution rate on our website's AI chat assistant.
The vendor dashboard showed our assistant with 0 resolutions out of 2 conversations. At our chat volume, that percentage was meaningless, and it pointed at the wrong problem. It looked like the AI was bad at answering. When I audited it this summer, the real issues were configuration: its instructions had been silently cut off at a character limit, mid-sentence, right where the rule for handing a customer to a human should have been. Part of its knowledge had been pulled from a staging copy of our site instead of the live one. The same dashboard also reported three active automated flows that were all switched off. None of that moves a resolution rate, and all of it hurts customers.
What replaced it: scripted test conversations graded against the truth. We send the assistant the kinds of questions our customers ask, like pricing for a specific home, a push for a ballpark number, an out-of-area request, and someone claiming to be the owner asking for wages and margins. Then we grade every answer against our actual sources: our pricing calculator, our live website, and our point-of-sale catalog. Checking what it tells customers against those sources is how we caught it saying our membership discounts were 10% and 15% when they're really 20% and 25%.
How we instrument it: no new software, just a fixed set of test prompts we can re-run whenever the setup changes, with each answer marked pass or fail against a source.
What changed: we stopped reading the scorecard. At Green Planet Cleaning Services, we'd rather check the actual work than trust a dashboard.
“A completion rate shows whether the system actually removed work.”
Aviad Faruz | Owner, FARUZO Jewelry
Read full expert response
I stopped using reply count as the success metric. An automated system can send many messages and still leave the difficult conversations for a person. In my venue WhatsApp workflow, I track the share of inquiries completed without human intervention. The workflow classifies each inquiry and selects a reviewed reply template from Google Sheets. It first ran in private note mode, where I could inspect drafts before anything reached a customer. It later handled 63% of incoming inquiries automatically, while unfamiliar cases still went to a person. That completion rate shows whether the system actually removed work.
At low volume, a percentage can be noise. At any volume, reply count can hide how much difficult work still reaches a person. Scripted conversations checked against real sources and a completion measure expose different failures.
So what’s the pattern?
Almost every replacement measures what happened after the answer: the customer came back, reordered, booked, used the product, reopened the issue, or moved to another channel. The other recurring ingredient is less glamorous. Someone tags and reads conversations by hand.
Nobody in this group began with a new analytics purchase. They began by joining chat to an outcome the business already cared about, then spending hours on definitions, tagging, and review.
If you are deciding what automation should be worth, start with the outcome in our chatbot ROI guide, test the assumptions in the ROI calculator, and build a review loop that can teach the chatbot from failed conversations.
Expert responses were provided through Connectively and are reproduced with permission. Figures and outcomes are each contributor’s own claims, not HoverBot findings.
Frequently asked questions
- Why can chatbot resolution rate be misleading?
- Chatbot resolution rate can count a conversation as solved when the customer simply stops replying. That makes a successful answer indistinguishable from frustration or abandonment. A stronger measure checks whether the same customer returns with the same issue, reopens the case, completes the intended task, or confirms that the problem was actually resolved.
- What should replace chatbot deflection rate?
- Deflection rate is more useful when paired with an outcome after the chat. Teams can track whether the customer contacts support again, calls another channel, completes a purchase, books a service, or uses the product. The replacement should reflect the task the customer wanted to finish, not merely the absence of a human conversation.
- How should support teams measure chatbot response time?
- Measure time to the first useful response, not time to an automated acknowledgment. A useful response either answers the question or asks for information the customer must provide. Teams can label messages as acknowledgments or substantive replies, then stop the timer only when the conversation genuinely moves toward resolution.
- Do better chatbot metrics require new analytics tools?
- Not necessarily. Several leaders in this roundup used existing helpdesk tags, order records, product events, test conversations, or simple spreadsheets. The recurring cost was human review: defining consistent tags, joining conversations to later actions, reading repeat contacts, and checking answers against approved sources. Better measurement often costs attention before it costs software.
Sources
About the author
AI Product Engineering Team
Cross-functional team of AI engineers, product managers, and support operators building customer-facing chatbot systems in production environments. We ship weekly releases informed by production telemetry, closed-loop conversation reviews, and benchmark-driven evaluation cycles.
- Customer support automation and intelligent routing systems
- RAG pipeline design and guardrails for regulated workflows
- Operational analytics and closed-loop quality improvement
- Multilingual NLP and entity-level PII masking pipelines
- Production deployments across e-commerce, real estate, and SaaS verticals


