Most free trials are wasted in the first ten minutes. Someone signs up, types "what are your opening hours?", gets a correct reply, types "tell me a joke", laughs, and closes the tab. A week later the trial ends and nobody has learned whether the thing can handle the enquiries the business actually gets.
An AI receptionist free trial is only useful if you treat it like an exam you set in advance. You pick the questions before you see the product, you decide what a pass looks like and you write the scores down. It takes an afternoon.
This plan is written for a website assistant, the kind that answers visitors in a chat window. It works on a free plan with a small message allowance, and it is honest about what that allowance cannot tell you.
Build the test set before you sign up
Open your email, contact form submissions and any notes from phone calls over the last two months. You are looking for the questions people really ask, in the words they really use, including the badly spelled ones.
Copy out about twenty. Aim for a spread like this:
- Ten routine questions your website already answers (prices, hours, service area, what is included).
- Four questions that need a follow-up before anyone could answer ("how much for a kitchen?").
- Four questions your content does not cover at all.
- Two messages from annoyed or confused customers.
The third group is the one people skip, and it is the most important. An assistant that invents an answer to a question you never taught it is a liability, however polished the rest looks. The guide on testing with real past enquiries goes further into picking the awkward ones.
Teach it the way you would for real
Do not paste in one perfect paragraph written for the test. Give it what you would actually give it on a normal Tuesday: your existing documents, your price list, a plain description of the business.
Be aware of what the free tier includes. On SpideyChat, for example, the free plan gives you one assistant and 100 messages a month, and the website crawler and lead capture sit on Starter. So on free you would teach by describing the business, uploading documents or adding Q&A pairs. That is enough to judge answer quality, which is the main thing a free trial should prove.
Run the questions and score every reply
Type each question as a visitor would, in a fresh conversation where possible. Score straight away, not at the end.
| Criterion | 2 points | 1 point | 0 points |
|---|---|---|---|
| Accuracy | Correct and complete | Correct but vague | Wrong or invented |
| Follow-up | Asked a sensible question when needed | Asked something unhelpful | Guessed instead of asking |
| Knows its limits | Said it did not know and offered a next step | Hedged | Made something up |
| Route to a person | Offered contact or handover when right | Only on request | Never |
| Tone | Sounds like your business | Neutral | Off-putting |
Not every criterion applies to every question. Score what applies and note the rest as "n/a".
Here is what a good result looks like on one of the uncovered questions:
Visitor: do you do work on listed buildings Assistant: I don't have information about listed building work, so I don't want to guess. If you leave your name and email, the team can confirm whether it's something they take on.
That reply would score 2 on knowing its limits and 2 on routing to a person. A reply that said "Yes, we have extensive experience with listed buildings" would score 0, even though it reads nicely.
Worked example: budgeting a small allowance
Say your free plan allows 100 messages a month. Treat this as an example and check how your vendor counts messages, because some count only visitor messages and others count both sides.
- 20 test questions, each averaging 3 visitor turns = 60 messages
- A second pass on the 6 questions that failed, after you fix the content = 18 messages
- Buffer for trying the welcome message and suggested prompts = 22 messages
That fits, just. If each side of the exchange counts, halve the test set to ten and pick the hardest ten. It is better to test fewer questions properly than to run out halfway through.
Reading the scores honestly
Add up the points per group, not just the total. The groups tell you different things.
- Routine questions scoring low usually means your content is thin or buried, not that the product is poor. Fix the content and re-run.
- Follow-up questions scoring low is a product behaviour. If it guesses at a kitchen price instead of asking about size, that will not improve much with more documents.
- Uncovered questions scoring low is the dealbreaker. Walk away if it invents answers, however well it did elsewhere.
- Annoyed customers scoring low is often fixable with a persona instruction and a handover rule.
What a free plan cannot prove
Be clear about the gaps, so you do not over-read a good score. A free test cannot show you how the assistant copes with a few hundred real visitors, how much time you will spend keeping the knowledge base current, or whether leads actually arrive in a form your team will follow up. It also cannot show you how the product behaves on your busiest day. The post on what a free plan can prove covers this from the budget side.
Your next move
Build the twenty-question set today, before you look at any product, and save it as a file. Then run it on each assistant you are considering, using the same sheet. If you want to see behaviour before creating an account, try the assistant with a few of your uncovered questions first.
If the scores are good, the sensible next step is a single low-risk page for a month on a paid tier, where lead capture and volume can be tested properly. Check pricing for the tier that includes what you need, and keep the test set: you will use it again every time you change the content.