When business owners test a new assistant, they almost always type the same thing: "What are your opening hours?" It answers correctly. They try "How much does it cost?" and it gives the price from the pricing page. Satisfied, they switch it on.
Then a real visitor types "hi do u do the thing where u come out and look at the damp b4 quoting, we're in the new estate off the ring road" and the answer is noticeably worse.
Questions we make up are tidy because we already know the answer and how the website phrases it. To test a chatbot with real questions, you need the ones customers actually sent, with the spelling, the context and the three things asked at once. You already have a month of them sitting in your inbox.
Where to find last month's enquiries
Spend twenty minutes gathering raw material. Look in every place questions arrive, because each channel attracts a slightly different kind.
- Email inbox: longer questions, often with background and several parts.
- Contact form submissions: short and direct, sometimes just a single line.
- Social media messages: casual wording, abbreviations, questions about availability.
- Notes from phone calls: if anyone jots down what callers asked, those count too, even though the assistant only works on your website.
- Existing chat history: if the assistant is already live, pick conversations where the answer was weak.
Aim for around thirty. Remove names, phone numbers, addresses and anything that identifies a person. Keep the wording otherwise, typos and all. That messiness is what you are testing.
Building the test set
Put the questions in a spreadsheet, one per row. Then check the spread, because thirty questions about price will tell you nothing about returns or service areas.
- Tag each question with a topic: pricing, availability, service area, how it works, problems or complaints, and anything specific to your trade.
- Look for topics with no questions at all. Add one or two real examples from further back if you can find them.
- Include at least three questions you would want the assistant to decline or hand over, such as a complaint, a request for someone's booking details or a question needing a qualified opinion.
- Include at least two questions whose answer is not on your website, so you can see how it handles gaps.
- Next to each question, write the correct answer in a few words, as your best staff member would give it.
That last column matters most. Without it, you end up judging answers by whether they sound confident, which is exactly the wrong test.
Scoring answers without overthinking it
Run each question through your test window, paste the reply into the spreadsheet, and score it. A simple rubric is enough to evaluate answers consistently, even across different people doing the scoring.
| Score | Accuracy | Completeness | Next step |
|---|---|---|---|
| 2 | Every fact matches your correct answer | Covers what was asked, including every part | Offers a sensible next action (details, ticket, handover) |
| 1 | Correct but vague, or slightly out of date | Answers some parts of a multi-part question | Next step is missing but not needed |
| 0 | Wrong, or says something you do not do | Misses the actual question | Leaves the visitor stuck |
One rule overrides the table: any invented fact, such as a price, a policy or a service that does not exist, fails the whole question regardless of the other scores. A well-grounded assistant should say it does not have the information and offer to take the enquiry. An answer like this one should score full marks for honesty even though it could not help:
Visitor: do u do the damp survey before quoting, we're on the new estate off the ring road Assistant: I don't have details on damp surveys or whether that estate is in our area, so I don't want to guess. If you leave your postcode and a number, the team will confirm both and get back to you.
And this one fails, however helpful it sounds:
Assistant: Yes, we offer a free damp survey across the whole area and can usually visit within 48 hours.
If nothing on your site mentions free surveys or a 48-hour visit, that reply has made a promise your team now has to deal with.
A worked example of reading the results
Here is how a first run might look for a hypothetical building repair firm. Swap in your own numbers.
Say you tested 30 questions. Six maximum per question, so 180 possible. The assistant scored 131, and two questions contained invented details.
- Headline score: 131 of 180 is about 73%. Useful as a benchmark to compare future runs against, not as a verdict.
- The two fails: both came from the same page, which used vague marketing language about "fast, free assessments". Rewrite that page.
- Low completeness on five questions: all were multi-part questions from email. Add a line to the persona asking it to address each part of a question in turn.
- Zero on next step for four questions: lead capture was switched off. Switch it on.
- Three knowledge gap answers: service area questions. Add a Q&A pair listing the towns you cover.
That is four changes, each tied to specific evidence. Most first runs look like this: a few content problems, one or two settings, and a handful of gaps. The guide on what to do when your agent keeps saying I don't know covers gap-filling in detail.
Re-running after every change
This is where most teams stop, and it is the step that makes the whole exercise worthwhile. Once you have made the fixes, run the entire set again, not just the questions that failed.
Changes have side effects. A new Q&A pair about service areas can start appearing in answers about delivery. A persona instruction to be more concise can make multi-part answers worse. A freshly crawled page can replace an accurate older answer with a vaguer one. You only catch these by re-running everything.
Keep each run as a new column in the spreadsheet with the date. Over a few months you get a simple record of whether the assistant is improving, and a reason for every change you made. Add new real questions monthly, especially from the unanswered questions list in your analytics, and retire any that no longer apply.
Start with this afternoon
Open your email and contact form, copy the last thirty genuine enquiries into a spreadsheet, and strip out personal details. Write your correct answer next to each. If you have not set up an assistant yet, a free account with your site crawled is enough to run the test, and the free trial test plan shows how to fit this into an evaluation period.
Score the first run honestly, fix the biggest problem first, and re-run the whole set before you change anything else. If you want help writing answers for the gaps you find, the FAQ generator is a quick way to draft Q&A pairs you can then correct.