Open most chatbot dashboards and you'll see a wall of numbers. Total conversations. Messages sent. Average session length. It looks impressive and tells you almost nothing about whether the thing is making you money or costing you customers. A bot can have thousands of conversations and still be quietly failing.
The fix isn't more data. It's the right five or six numbers, watched over time, each tied to a decision you'd actually make. Here's the shortlist for an online store, and what each one is really telling you.
Resolution rate: is the bot absorbing the load
The first question is simple. Of all the conversations your bot starts, how many does it finish without dragging a human in? That's your resolution rate, sometimes called containment.
A high number means the bot is doing its main job, soaking up the routine questions so your team doesn't have to. But read it carefully, because resolution rate is easy to fake. A bot that says "I can't help with that, try our contact page" technically ended the conversation, and that's not a resolution, that's a dead end. Always pair this number with satisfaction, below, so you know the conversations ended well and not just quietly.
There's a sharper cut of this number worth pulling: resolution rate by topic. A bot might resolve shipping questions cleanly and fumble every sizing question. The blended average hides that. Break it down and you find exactly which piece of content to fix next, instead of guessing.
The revenue numbers: conversion and order value
For a store, deflecting support tickets is only half the value. The other half is revenue, and this is where most people forget to look.
Start with conversation-to-sale. Track how many chatbot conversations end in a purchase, and compare shoppers who chatted with those who didn't. If people who talk to the bot buy more often, the bot is doing sales work, not just support. This is the metric that turns a chatbot from a cost center into something you'd happily pay more for.
Take Fernway, a fictional houseplant shop. They noticed that visitors who asked the bot a care question before buying converted at a noticeably higher rate than those who didn't, because the bot answered the exact worry holding them back ("will this survive in low light"). That single insight told them the bot was worth expanding, not trimming.
Then look at how much people spend, not just whether they buy. Two numbers matter here: the revenue tied to conversations that touched the bot, and the average order value of chatters versus non-chatters. A good recommendation nudges people toward the right item, and sometimes a better or larger one. If your bot suggests a complementary product ("that planter needs a drainage tray, want me to add one?") and shoppers take it, average order value climbs. Watching this tells you whether your recommendation logic is pulling its weight or just being polite.
Escalation rate: your customers' to-do list
Escalation rate is the flip side of resolution. What share of conversations does the bot hand to a human, and, more usefully, why?
The raw number matters less than the reasons behind it. If the bot keeps escalating the same three question types, you've found your next improvement list. Maybe it doesn't know your international shipping rules, or it fumbles sizing. Feed it that content and the escalations for those topics drop. Escalation data is basically a to-do list your customers write for you, sorted by how often each gap actually comes up.
A rising escalation rate isn't automatically bad, either. If it climbs because traffic grew or because you added a genuinely complex product line, that's context, not failure. Read it next to volume before you panic about it.
Satisfaction: the guardrail on the rest
Every number above can look great while customers quietly hate the experience. Satisfaction is the guardrail.
A quick thumbs up or down after a conversation, or a one-tap rating, is enough. You're not after a perfect score, you're after a trend and a warning system. A resolution rate that climbs while satisfaction falls is a red flag: the bot is ending conversations, but not happily. Watched together, these two numbers keep each other honest.
Don't over-engineer the collection, though. A simple optional rating that most people skip still surfaces the trend you need, and it won't annoy the shoppers who just wanted their answer and to get on with their day.
The vanity numbers to mostly ignore
Some metrics feel important and aren't. Total conversation count goes up when traffic goes up and tells you nothing about quality. Average messages per chat can mean the bot is thorough or that it's confusing people into asking again. Session length is the same trap: long could be engaged or lost.
None of these are useless, but none should drive a decision on their own. Treat them as texture, not headlines. If total conversations doubled this month, that's a prompt to go ask why, not a result to celebrate by itself.
Turning numbers into changes
Metrics only pay off when they change what you do. Here's the shortlist again, side by side with the trap to avoid for each:
| Metric | What it answers | Watch out for |
|---|---|---|
| Resolution rate | Is the bot handling routine load? | Dead-end "endings" that aren't real answers |
| Conversation-to-sale | Is the bot driving purchases? | Judging the bot on support alone |
| Influenced revenue / AOV | Is each chatter worth more? | Recommendations that don't fit the shopper |
| Escalation rate | Where does the bot hit limits? | Chasing the number instead of the reasons |
| Satisfaction | Are people glad they chatted? | A rising resolution rate hiding falling happiness |
A simple loop keeps them honest:
- Pick your handful of real metrics and check them weekly, not obsessively.
- Find the one that's weakest relative to where you want it.
- Read ten real transcripts behind that number to see what's actually happening.
- Make one change: better content, a new answer, a smoother recommendation.
- Watch the metric for a couple of weeks before you touch anything else.
That last step is where discipline pays. Change five things at once and you'll never know which one helped. Change one, and the metric tells you the truth.
Most hosted tools surface these numbers for you. In SpideyChat you can see resolution and escalation alongside the conversations that led to leads or sales, so the "what" and the "why" sit in the same place and you're not stitching together three dashboards to answer one question.
Start smaller than you think. Pick just two numbers this week, resolution rate and satisfaction, and read the transcripts behind them. That pairing alone will tell you more about whether your bot is earning its place than a dashboard with thirty tiles ever could.