Why this checklist exists, and why we included the questions that hurt us
Every AI receptionist demo you will sit through this year is going to be impressive. The voice will sound natural, the sample call will go perfectly, and the appointment will appear on a screen at exactly the right moment. That is what demos are engineered to do, and it tells you almost nothing about how the system behaves on a Tuesday at 6:40 PM when a caller with a thick accent and a dying phone battery is trying to describe a problem the AI has never encountered.
As of August 2026 this category has enough vendors in it that the honest differentiator is no longer whether the AI can hold a conversation — most can. It is what happens at the edges: the calls it cannot handle, the data it cannot write, the invoice line you did not expect, and the exit you did not ask about. This is a buyer's checklist for those edges, organized into nine areas with the specific questions to ask, what a good answer sounds like, and what a bad answer is trying to avoid saying.
We publish it knowing it will be used against us. Several questions below are ones where Run with Jarvis is not the automatic winner, and we have flagged those explicitly rather than quietly skipping them. A checklist that only asks questions the author passes is marketing, not diligence.
1. Pricing transparency
Start here, because a vendor's posture on pricing predicts their posture on everything else.
The questions
- What is the monthly price, published, for the tier I would actually buy?
- Is pricing per-minute, per-call, per-seat, or a flat plan?
- What exactly counts as a billable minute — ring time, hold time, voicemail, transferred calls, calls that connect and immediately hang up?
- What is the overage rate, in writing?
- Is there a setup, onboarding, or implementation fee?
- What is the contract length, and what is the cancellation notice period?
- Does the price change after a promotional period?
- What is not included that I will end up paying for — integrations, extra numbers, additional languages, recording storage?
What a good answer sounds like
A number, on a public page, that matches what the salesperson says. "Contact us for pricing" in this category almost always means the price is set by what the salesperson thinks you will pay. That is not automatically dishonest, but it does mean your quote is not comparable to anyone else's, which defeats the entire purpose of an evaluation.
Watch the billing unit closely. Per-call pricing punishes you for the exact behavior you want — answering everything, including short calls and wrong numbers. Per-seat pricing punishes you for growing. Per-minute-only pricing with no included pool makes your bill unpredictable in your busiest month, which is when cash is already tight.
For reference, our numbers are published and flat: $500/mo for Core with 500 AI call minutes and $0.45/min overage, $750/mo for Pro with 1,000 minutes and $0.40/min overage, $1,200/mo for Elite with 2,500 minutes and $0.35/min overage. Zero setup fees, unlimited users, no per-call fees, month-to-month on every tier. All of it lives on the pricing page, and the underlying math is worked out example-by-example in what AI operations actually cost.
Push back on us here
Ask us the uncomfortable version: $500/mo is a real monthly commitment for a one-truck operation, and there is no cheaper entry tier. If your business takes 25 calls a month, the honest answer is that the math is harder to justify than it is at 150 calls, and you should make us show the arithmetic against your actual missed-call count rather than a category average. Also ask what happens to your bill in your peak month — overage is real, and you should see the worst-case number before you sign, not after.
2. What happens on a call it cannot handle
This is the single most important section in this checklist, and it is the one buyers skip.
Every AI receptionist will hit calls it cannot complete. An unusual piece of equipment. A caller who wants to argue about a previous invoice. A commercial account with negotiated pricing. A situation that needs judgment. The question is not whether that happens — it is what the system does when it does.
The questions
- When the AI cannot resolve a call, what are the possible endings? There are only three: live transfer, structured callback capture, or a dead end.
- If it transfers to a human and nobody answers, what then? This is the critical follow-up.
- Is the transfer a blind hand-off, or can the AI come back and capture details if the transfer fails?
- Does the escalation notify anyone actively — a text, an app notification, a team channel — or does it just sit in a log?
- Can I hear a recording of a real failed call, not a scripted demo?
- Can I define my own escalation rules by call type, keyword, time of day, or customer?
What a good answer sounds like
The right architecture is capture first, transfer second. The AI collects the caller's name, number, vehicle or property, and the nature of the problem before attempting a hand-off, so that if the transfer fails, the lead still exists. Vendors who transfer first and capture nothing are betting your revenue on somebody picking up.
The follow-up question about failed transfers matters more than it sounds. Many phone systems perform what amounts to a one-way hand-off: once the transfer fires, the AI is gone and cannot take over again. If the destination rings out to a full voicemail box, the caller lands in a dead end with nothing recorded. Ask specifically: "if the human line does not answer, does the AI still have the lead?" If the answer is vague, assume the answer is no.
And insist on question 13. A vendor who will not play you a call that went badly either does not review their own failures or does not want you to hear them.
Push back on us here
Our own answer is capture-first with an active dispatch notification, and we will show you failed-call recordings. But the honest caveat is this: an escalation is only as good as the person who reads it. If nobody on your team watches the alert channel, a perfect capture-first architecture still ends in a customer who never got called back. That is not a vendor problem you can buy your way out of — ask us how we would set the notification path up for your team, and then hold your own team to it.
3. Can it actually write, or only read?
The dividing line in this category is write access. A system that can read your calendar to check availability and then email you a request is fundamentally a message service. A system that creates the appointment is an employee.
The questions
- Can it create an appointment in my calendar, live, while I watch?
- Does that appointment carry the right duration, address, service type, and assigned technician — or just a time and a name?
- Can it reschedule and cancel, not just create?
- Does it create or update the customer record in the CRM, with call history attached?
- Does it check travel time and service-area boundaries before offering a slot?
- What does it do about double-booking and conflicts?
- Can it quote from my actual price list, and how do I load and update that list?
What a good answer sounds like
Make them prove it in a live demo using your calendar or a copy of it. Watch the appointment appear. Then ask them to reschedule it by phone. Rescheduling is the capability that reveals whether the integration is real — creating a record is easy, modifying an existing one requires the system to find it first, which requires genuine two-way access.
The quoting question deserves its own scrutiny. There is a large gap between an AI that reads a static price range off a script and one that looks up a specific service in your catalog, applies your trip fee and any after-hours surcharge, and speaks one all-in number. Ask how prices get updated, how long that takes, and what the AI says when a job is not in the catalog — the correct behavior is "a technician will confirm your exact setup," not an invented figure. The mechanics are covered in our guide to quoting prices from a price list.
Then check where the lead lands. If the AI books but the customer record lives in a separate system you have to reconcile by hand, you have bought half a solution. Our piece on an AI receptionist with built-in CRM and lead tracking explains why the seam between answering and record-keeping is where most stacks leak.
4. Language coverage
The questions
- Which languages does it handle, and on which number?
- Does the caller have to choose a language, or is it detected?
- If a call starts in one language and switches, does the conversation keep its context — the vehicle, the address, the quote — or does the caller start over?
- Are the non-English prompts written by native speakers or machine-translated?
- Do bilingual calls cost more?
What a good answer sounds like
Same number, automatic detection, and — this is the one to press — context preserved across a language switch. Some systems handle a language change by handing the call to a separate assistant, which in many implementations means the second assistant starts blind: no vehicle, no address, no quote, forcing the caller to repeat everything they just said. That is a real failure mode, not a hypothetical, and it produces exactly the frustrated hang-up you bought the system to prevent.
Ask for a recorded bilingual call. If the vendor only has English recordings, their Spanish coverage is a checkbox rather than a product.
5. Data ownership, export, and the exit
Nobody negotiates the exit while they are excited about the entrance. Do it anyway, because this is where switching costs are manufactured.
The questions
- Who owns the call recordings, transcripts, contacts, and appointment history?
- Can I export all of it, on demand, in a standard format, without a fee?
- How long is my data retained after I cancel, and is there a window where I can still pull it?
- Are my recordings and transcripts used to train models that serve other customers?
- If I leave, do I keep my phone number?
What a good answer sounds like
"You own it, you can export it any time, here is the format." Anything softer is a lock-in mechanism. Question 31 is the sleeper. If the vendor provisioned your tracking numbers, ask explicitly whether those numbers port out. A number printed on your trucks, your yard signs, and every directory listing you have ever claimed is not a small asset — a vendor who controls it controls your ability to leave. Get portability in writing.
Question 30 is increasingly worth asking directly. It is a reasonable business practice for a vendor to improve their system using aggregate data, and it is also reasonable for you to know whether your customers' recorded conversations are part of that. The answer should be specific, not reassuring.
6. Recording, consent, and compliance
The questions
- Does it record calls by default, and can I turn that off per number or per state?
- How is consent handled — a disclosure at the top of the call, and is that disclosure configurable?
- Where are recordings stored, for how long, and who at the vendor can access them?
- How are text messages handled from a compliance standpoint — opt-out keywords, message classification, sender registration?
What a good answer sounds like
Recording rules vary meaningfully by state, and consent requirements are not a detail you can retrofit after a complaint. A competent vendor will already have a configurable disclosure and will be able to tell you how it behaves for two-party-consent states without you raising it first. Our call recording consent and compliance guide covers the operational side; the underlying regulatory context sits with the FCC (fcc.gov) for telephone practices and the FTC (ftc.gov) for consumer-protection and telemarketing rules. Neither the vendor nor this checklist is a substitute for your own counsel on the specifics of your states.
On texting, ask about sender registration and automatic opt-out handling. A vendor whose SMS is not properly registered will see deliverability quietly degrade, and you will experience that as customers who "never got the text."
7. Integration reality
"Integrates with QuickBooks" is a phrase, not a specification.
The questions
- Which specific systems does it write to, and is the sync one-way or two-way?
- What syncs to accounting — customers, invoices, payments, all three — and how often?
- Does it know my price list, and where does that list live?
- What happens to the sync when a record conflicts on both sides?
- Is the integration included in my plan, or a paid add-on?
What a good answer sounds like
Ask for the field list. A real accounting integration syncs customers, invoices, and payments bidirectionally on a defined schedule, and the vendor can tell you which fields map where. A shallow one pushes a summary and leaves you rekeying. Our QuickBooks sync guide sets out what a complete sync actually covers, which makes it easy to spot a partial one.
Then step back and ask the structural question: how many vendors are you assembling? If the answering AI is from one company, the calendar from another, the CRM from a third, and call tracking from a fourth, every integration in that chain is a place where a booked job can fail to become an invoiced one — and every one of those seams is your operational responsibility, not theirs. We wrote about that tradeoff honestly in all-in-one versus point solutions, including the cases where separate best-in-class tools genuinely win.
Push back on us here
The all-in-one argument cuts both ways, and you should say so out loud: buying answering, booking, CRM, POS, and dispatch from one vendor concentrates risk. If we have a bad day, more of your operation is affected than if you ran four separate tools. Our counter is that four tools have four times the failure surface and no single owner when something breaks — but that is a genuine tradeoff, not a slam dunk, and any vendor telling you their consolidation has no downside is not being straight with you.
8. Uptime, fallback, and what happens on a bad day
Voice AI depends on a chain of services — telephony carrier, speech recognition, a language model, text-to-speech, your calendar. Any link can have a bad minute.
The questions
- What is the published uptime record, and is there a status page?
- When the platform is unavailable, what does a caller hear?
- Is there a configured fallback that routes callers to a human line instead of silence?
- How am I notified of an incident, and how fast?
- What happens if my account hits a billing or credit problem — do calls stop, and do I get warned first?
What a good answer sounds like
The right answer is a defined fallback path: when the AI cannot answer, the carrier is configured to forward to a live number rather than play an error message. Ask whether that fallback is set up by default or something you have to request. A vendor who has genuinely thought about outages will describe this without prompting; one who has not will talk about redundancy in the abstract.
The billing question is not paranoia. Usage-based platforms can suspend service when a balance runs dry, and the failure mode is callers hearing an error instead of a receptionist. Ask whether balance alerts and auto-recharge exist, and turn them on.
Push back on us here
Ask us for our incident history and our fallback configuration specifically, and ask whether it is enabled on your numbers on day one or left as a setup task. "We have high uptime" is not an answer. "Here is the number your calls forward to if we are down, and here is how we verify it" is.
9. The trial that actually proves something
Most trials are theater. Here is how to run one that generates evidence.
Design rules
- Minimum two weeks, covering two full weekends and several evenings. A weekday-only trial never tests the hours where the coverage gap is worst, which is exactly where an answering layer earns most of its money.
- Real calls on a real tracked number. A sandbox tests the software; live traffic tests the fit. Route a portion of your actual inbound volume through it.
- Load your real price list first. A trial with generic pricing proves nothing about quoting, which is the highest-value capability.
- Do not tell your team it is a trial on the escalation side. You want to see how escalations actually get handled, not how they get handled when everyone is watching.
The five things to measure
- Answer speed and abandon rate. Every call answered on the first ring, no exceptions? Pull the call detail records and verify rather than trusting a dashboard.
- Quote accuracy. Sample 20 calls where a price was spoken and check each against your price list, including the trip fee and any after-hours surcharge. Any mid-call price revision is a failure.
- Booking correctness. For every appointment created, verify the date, the time, the duration, the address, and the assigned tech. The date and time are the ones to check hardest — a system that resolves "tomorrow at 2" into the wrong slot will generate no-shows and angry customers, and it is a surprisingly common defect.
- Escalation completion. Count the escalations and count how many reached a human within your target window. This is the number that predicts whether the system will lose you a customer.
- Bilingual handling, if it applies to your market — including a call that switches language mid-conversation.
The one question to ask at the end of the trial
Not "did it work" but "what did it do that I did not expect?" Every AI answering system develops behaviors that were not in the demo. Some are pleasant surprises; some are the thing that quietly costs you a customer in month four. Two weeks of real traffic is enough to surface a few, and a vendor who reacts well when you bring them one is worth more than a vendor whose demo was smoother.
A shortcut for the tier decision
Once you have chosen a vendor, sizing is a separate exercise and much simpler: it comes down to your monthly call volume, whether you spend on paid channels, and whether you want a marketing layer. Here is the compact version for Run with Jarvis, and the full walkthrough is in how to choose an AI receptionist plan — which is about picking a tier within our platform, a different question from the vendor-selection one this checklist covers.
| Question | Points to Core ($500) | Points to Pro ($750) | Points to Elite ($1,200) |
|---|---|---|---|
| Monthly AI-answered calls | Up to ~125 | ~125-250 | ~250+ |
| Included AI minutes | 500 | 1,000 | 2,500 |
| Overage rate | $0.45/min | $0.40/min | $0.35/min |
| Do you spend on Google Ads or LSA? | No | Yes | Yes, heavily |
| Need call recording and transcripts? | Not required | Yes | Yes |
| Want AI campaign and ad management? | No | No | Yes |
| Want natural-language operations by text? | No | No | Yes |
| Operations core (answering, booking, CRM, POS, dispatch, invoicing) | Included | Included | Included |
The row to notice is the last one: the complete operations system is in the first tier. You move up for measurement and marketing, not to unlock invoicing or dispatch. If any vendor's tier structure puts the basics behind a middle plan, that is a pricing question worth raising before a product question.
The Elite row about natural-language operations refers to the Jarvis brain — asking your business questions over SMS or WhatsApp instead of opening dashboards. Worth a specific evaluation question of its own: ask any vendor offering an "AI assistant" what it can actually do versus what it can only report on, because the gap between the two is usually wide.
The one-page version
If you take nothing else from this, take these six:
- Get the price in writing, including the overage rate and what counts as a billable minute.
- Ask what happens when the AI fails, and what happens when the human transfer also fails.
- Make them create, then reschedule, an appointment in your calendar while you watch.
- Confirm you own and can export your data — and that you keep your phone number.
- Ask what a caller hears when the vendor's platform is down.
- Run two real weeks including weekends, with your real price list loaded.
Any vendor who answers all six cleanly is worth serious consideration. Any vendor who deflects on more than one is telling you something. And if you want to run this list at us directly, get in touch — we would rather be evaluated on it than talked past it.



