Best Practices· 5 min read

Test Your Assistant With Real Enquiries From Last Month

Build a test set from real emails and messages, score each answer against a simple rubric, and re-run it every time you change your assistant.


When business owners test a new assistant, they almost always type the same thing: "What are your opening hours?" It answers correctly. They try "How much does it cost?" and it gives the price from the pricing page. Satisfied, they switch it on.

Then a real visitor types "hi do u do the thing where u come out and look at the damp b4 quoting, we're in the new estate off the ring road" and the answer is noticeably worse.

Questions we make up are tidy because we already know the answer and how the website phrases it. To test a chatbot with real questions, you need the ones customers actually sent, with the spelling, the context and the three things asked at once. You already have a month of them sitting in your inbox.

Where to find last month's enquiries

Spend twenty minutes gathering raw material. Look in every place questions arrive, because each channel attracts a slightly different kind.

Aim for around thirty. Remove names, phone numbers, addresses and anything that identifies a person. Keep the wording otherwise, typos and all. That messiness is what you are testing.

Building the test set

Put the questions in a spreadsheet, one per row. Then check the spread, because thirty questions about price will tell you nothing about returns or service areas.

  1. Tag each question with a topic: pricing, availability, service area, how it works, problems or complaints, and anything specific to your trade.
  2. Look for topics with no questions at all. Add one or two real examples from further back if you can find them.
  3. Include at least three questions you would want the assistant to decline or hand over, such as a complaint, a request for someone's booking details or a question needing a qualified opinion.
  4. Include at least two questions whose answer is not on your website, so you can see how it handles gaps.
  5. Next to each question, write the correct answer in a few words, as your best staff member would give it.

That last column matters most. Without it, you end up judging answers by whether they sound confident, which is exactly the wrong test.

Scoring answers without overthinking it

Run each question through your test window, paste the reply into the spreadsheet, and score it. A simple rubric is enough to evaluate answers consistently, even across different people doing the scoring.

Score Accuracy Completeness Next step
2 Every fact matches your correct answer Covers what was asked, including every part Offers a sensible next action (details, ticket, handover)
1 Correct but vague, or slightly out of date Answers some parts of a multi-part question Next step is missing but not needed
0 Wrong, or says something you do not do Misses the actual question Leaves the visitor stuck

One rule overrides the table: any invented fact, such as a price, a policy or a service that does not exist, fails the whole question regardless of the other scores. A well-grounded assistant should say it does not have the information and offer to take the enquiry. An answer like this one should score full marks for honesty even though it could not help:

Visitor: do u do the damp survey before quoting, we're on the new estate off the ring road Assistant: I don't have details on damp surveys or whether that estate is in our area, so I don't want to guess. If you leave your postcode and a number, the team will confirm both and get back to you.

And this one fails, however helpful it sounds:

Assistant: Yes, we offer a free damp survey across the whole area and can usually visit within 48 hours.

If nothing on your site mentions free surveys or a 48-hour visit, that reply has made a promise your team now has to deal with.

A worked example of reading the results

Here is how a first run might look for a hypothetical building repair firm. Swap in your own numbers.

Say you tested 30 questions. Six maximum per question, so 180 possible. The assistant scored 131, and two questions contained invented details.

That is four changes, each tied to specific evidence. Most first runs look like this: a few content problems, one or two settings, and a handful of gaps. The guide on what to do when your agent keeps saying I don't know covers gap-filling in detail.

Re-running after every change

This is where most teams stop, and it is the step that makes the whole exercise worthwhile. Once you have made the fixes, run the entire set again, not just the questions that failed.

Changes have side effects. A new Q&A pair about service areas can start appearing in answers about delivery. A persona instruction to be more concise can make multi-part answers worse. A freshly crawled page can replace an accurate older answer with a vaguer one. You only catch these by re-running everything.

Keep each run as a new column in the spreadsheet with the date. Over a few months you get a simple record of whether the assistant is improving, and a reason for every change you made. Add new real questions monthly, especially from the unanswered questions list in your analytics, and retire any that no longer apply.

Start with this afternoon

Open your email and contact form, copy the last thirty genuine enquiries into a spreadsheet, and strip out personal details. Write your correct answer next to each. If you have not set up an assistant yet, a free account with your site crawled is enough to run the test, and the free trial test plan shows how to fit this into an evaluation period.

Score the first run honestly, fix the biggest problem first, and re-run the whole set before you change anything else. If you want help writing answers for the gaps you find, the FAQ generator is a quick way to draft Q&A pairs you can then correct.

Frequently asked questions

How many questions do I need to test a chatbot properly?
For a small business, 25 to 40 real questions usually covers the common topics and a few awkward ones. More is better only if it adds variety, not duplicates.
Is it all right to paste customer emails into a chatbot for testing?
Strip out names, contact details and anything personal first. You only need the question itself, reworded if necessary.
What score should an assistant reach before going live?
Aim for no invented facts at all and most answers fully correct. Partial answers that capture the enquiry properly are acceptable for launch while you fill content gaps.
Can someone who is not technical run these tests?
Yes. It needs a spreadsheet, a test window and someone who knows the business well enough to judge whether an answer is right.

Keep reading

Test Your Assistant With Real Enquiries From Last Month · SpideyChat