"It's getting a ton of conversations." That's the answer most people give when you ask how their chatbot is doing. It's also not an answer. A chatbot can be busy and useless at the same time, fielding hundreds of chats while resolving almost none of them and quietly annoying everyone who tries it.
Measuring a chatbot well means ignoring the numbers that flatter you and watching the ones that tell you whether customers actually got helped. There aren't many that matter, which is good news. You can track the real ones in a spreadsheet.
Vanity metrics versus the ones that pay rent
Total conversations, messages sent, "engagement" — these feel like progress and mean almost nothing on their own. A bot that gets a thousand chats and resolves none is worse than one that gets a hundred and resolves ninety. Volume tells you people are trying it. It says nothing about whether it worked.
The useful metrics all answer one question: did the customer get what they came for, and at what cost to your team? Everything else is decoration. If a number doesn't connect to a resolved question, a captured lead, or a saved hour, don't build your judgment on it.
The four numbers that actually matter
You can measure a chatbot's real value with four things.
- Resolution rate. Of all conversations, what share did the bot handle completely without a human stepping in and without the customer leaving frustrated? This is the closest thing to a headline number.
- Handoff quality. When the bot escalates, does the human receive the full context, or do they start from scratch? A clean handoff saves the interaction; a cold one wastes it.
- Lead capture. If the bot is meant to collect leads, how many qualified ones did it capture, and did they reach the right person?
- Why unresolved chats failed. For the conversations the bot couldn't handle, what was the reason? Missing content, a genuinely human-only question, or a bug?
That fourth one isn't a number; it's a habit; it's also the most valuable, because it turns every failure into a fix.
Read the transcripts, not just the dashboard
Dashboards tell you what happened. Transcripts tell you why. If you only do one thing to improve your chatbot, spend fifteen minutes a week reading actual conversations, especially the ones that ended badly.
You'll learn things no chart surfaces. The odd phrasing real customers use. The question you never thought to add to your content. The moment the bot confidently gave a slightly-wrong answer. The place where it should have handed off and didn't. Patterns jump out fast: if five people in a week asked about weekend hours and the bot fumbled each time, you've found a content gap worth thirty seconds to fix.
Pay special attention to the last message a customer sends before they leave without an answer. That final message is where the bot lost them, and it's the single most useful thing to read. Sometimes it's a question you never covered. Sometimes it's a sign of frustration ("just give me a person") that the bot ignored. Either way, that exit point tells you precisely what to change next, in a way no aggregate percentage can.
A simple scorecard
You don't need fancy tooling to start. A weekly line in a sheet is enough to see the trend.
| Week | Conversations | Resolved by bot | Escalated with context | Leads captured | Top failure reason |
|---|---|---|---|---|---|
| 1 | 120 | 58 | 22 | 9 | Missing shipping info |
| 2 | 135 | 79 | 25 | 12 | Sizing questions |
| 3 | 141 | 96 | 20 | 15 | (mostly resolved) |
The story here isn't any single cell; it's the climb in "resolved by bot" as failure reasons get fixed one by one. Week 1's shipping gap gets patched, so week 2's problem is sizing, which then gets patched too. That's what a healthy chatbot looks like on paper: a moving target of shrinking problems.
What "good" looks like at each stage
Resist the urge to chase a magic benchmark. What's good depends on your content quality and the mix of questions you get. A bot answering simple factual questions on a clean help center will resolve far more than one fielding messy, account-specific requests.
Early on, many teams see roughly half of conversations fully resolved, with the rest either escalating or trailing off. That's a fine starting point, not a failure. The number that matters is the direction. If resolution climbs week over week as you close content gaps, the bot is working and getting better. If it's flat and your failure reasons keep repeating, something's stuck and it's usually your source content, not the technology.
It also helps to segment. A blended resolution rate can hide a lot. Maybe the bot crushes shipping and hours questions while whiffing on anything about billing. The average looks fine, but a whole category is failing. Break your resolution number down by question type and the weak spots stop hiding behind the strong ones. That's often where the real work is: not lifting a mediocre bot everywhere at once, but rescuing the one or two topics dragging the whole thing down.
There's a business number worth watching too, even roughly. If the bot resolves, say, a hundred questions a week that a person would otherwise have handled, and each of those took a few minutes of someone's time, you can estimate the hours it's giving back. Don't overstate it, but a rough hours-saved figure is what turns "the bot seems useful" into a number your boss can weigh against its cost.
Turn what you find into changes
Metrics only matter if they drive action. Each failure reason maps to a specific move. Missing content means write the missing answer and add it to what the bot draws from. A question that should always reach a person means tighten your handoff rules so it routes cleanly. A bot answering confidently but wrong means your source material is unclear or contradictory, so fix the underlying page.
Take Loft & Larder, a small kitchenware shop. Their first month, the bot resolved a little over half of chats, and the top failure reason was people asking whether items were dishwasher-safe. That fact was missing from most product pages. They added it, the dishwasher questions vanished from the failure list, and resolution jumped the following week. In SpideyChat you'd spot that pattern in the conversation logs and update the source content directly, so the fix sticks. No new technology, just a gap closed.
Give any new chatbot two to four weeks of real traffic before you judge it, and treat those early numbers as a map of what to fix rather than a verdict. Watch resolution climb, keep reading the transcripts, and let the failure reasons tell you what to do next. A chatbot that's measured this way rarely stalls, because every week you're handing it the answer to whatever tripped it up last.