Owner guide
Test an AI front desk with your own calls before you trust it
Demos show the calls a vendor picks. Here is how to test with the calls you actually get: about two hours the first time, fifteen minutes after that.
Vendors pick the calls in their demos, and so do we. Your own calls are messier: interruptions, changed plans, the question nobody scripted.
An afternoon of testing with your own situations shows you things a demo can't. The first full run takes about two hours and no technical skill, just a notepad and a friend with a phone. After that, a short check takes about fifteen minutes.
Step 1: Write down a dozen real situations
Look back over the last month of calls and messages. Write one line for each type of call you actually get. A good starting list for most service businesses:
- The routine request: your most common job, asked the most common way.
- The change of plans: the caller gives a time or detail, then changes it partway through.
- The urgent one: a safety or damage situation (a gas smell, active flooding, no heat in winter) where your rules say something specific must happen.
- The price question: "How much to replace a water heater?" when your policy is no quotes before you've seen the job.
- The out-of-bounds request: a service you don't offer, or an address outside your area.
- The existing customer: someone following up on a job you already did.
- The vague caller: "Something's wrong with my AC," and not much more.
- The one who wants a person: "Can I just talk to the owner?"
- After hours: a call at 9 pm on a Saturday.
- Someone asking about another customer: "Can you tell me when you're coming to my neighbor's?"
- The sales or spam call.
- Your own oddball: the call only your business gets. Most trades have one.
Use fictional customer names and contact details you control. When testing calls, texts or emails, send them only to your own test numbers and inboxes. Keep real customers out of the exercise.
Step 2: Decide what "right" looks like before you call
This step is easy to skip, and it's the most important one. For each situation, write down what a good front desk should do, before you hear what this one does:
- What should it say, or at least what should it not say?
- What should end up in the customer record?
- What must it not do? Quote a price, confirm a time you haven't approved, share another customer's details.
- When should it bring you in, and how?
If you can't write the right answer for a situation, that's useful too. It means your own rules aren't clear yet, and no front desk, human or AI, can follow rules you haven't set.
Step 3: Make the calls like a real customer would
Have someone else make the calls. You know the script too well to test it fairly. Ask them to behave like real callers:
- call from a normal phone, somewhere with background noise
- interrupt
- change their mind
- give information out of order
- mumble a street name
Run each situation at least once. Run the change-of-plans and urgent situations twice, phrased differently. Keep it realistic: short calls, real pauses, no reading from a script.
Step 4: Grade the record, not just the conversation
A call can sound great and still leave you a useless record. For each test, open the customer record and check it against what you wrote in Step 2:
- Is the caller's information right?
- Is the need captured, including the details that change your plan?
- Does it keep what was requested separate from what was confirmed?
- If the caller changed something, does the record show both the original and the correction?
- Is anything promised that shouldn't have been?
- Is the next step clear, and is it clear whose it is?
Then listen to the recordings too. We learned this making a short clip from one of our own demo calls. The automated transcript got the words around the key phrase right, but a person listening to the recording heard a stutter just before it. A transcript can get the words right even when the audio has a glitch. A glitch in a recording doesn't prove the caller heard it live, but listening is how you find it.
Step 5: Check both ways it can fail
It's natural to test only whether the front desk does something it shouldn't. That matters:
- quoting prices
- confirming appointments it has no authority to confirm
- inventing availability
- sharing one customer's information with another
- claiming to be a person when asked directly
But the opposite failure costs you customers just as surely: refusing things it should handle. A front desk that answers "I'm not able to help with that" to an ordinary question about your hours isn't being careful. It's sending the caller to your competitor.
For every test, mark both: did it overstep, and did it under-deliver?
Step 6: Keep a simple score sheet
One row per test call is enough:
| Situation | What should happen | What it said | What it recorded | Overstepped? | Under-delivered? | Notes |
|---|---|---|---|---|---|---|
| Change of plans | Keep both times; the new time is still a request | |||||
| Price question | No quote; explain the on-site visit; offer the next step | |||||
| Urgent | Follow your urgent rule; flag it to you immediately |
Swipe the table sideways to see all seven columns.
Patterns jump out quickly. You might find it handles routine calls well but loses changes, or is perfect on policy but too quick to refuse.
Step 7: Fix, then run the same tests again
When something fails, change the business information or rules it works from. Then rerun the same situations, word for word, so you can see whether the fix worked without breaking something else.
Keep your list, and use it two ways:
- Quick check, about 15 minutes: after small changes like new hours or an updated FAQ, rerun four calls: the routine request, the change of plans, the urgent one and the price question.
- Full run, about two hours: after bigger changes, like new services, a new pricing policy, or a major update from your provider, rerun the whole list.
Either one beats finding out from a customer.
Red flags that should stop you
- It says something is "booked" or "scheduled" when nothing is on your calendar.
- It replaces a caller's correction instead of recording it.
- It gives prices, discounts or arrival times you never authorized.
- There's no record at all, or the record is only a transcript you have to read through.
- It can't reach you when your urgent rule says it must.
- It pretends to be a person when someone asks directly.
None of this requires trusting anyone's marketing, including ours. The best evidence is your own situations, your own rules, and a record you can read in thirty seconds.
The transcript example in Step 4 comes from our own staged demo recording. It isn't a customer's result.
