Testing an AI Agent Before Launch: 10 Must-Run Checks

How do you know the agent is ready to talk to a paying customer? Most owners answer the same way: open the chat, ask five or six questions they would ask themselves, get answers that sound excellent, approve go-live. That tests one thing only, that the agent can answer someone who already knows what they want. Week-one customers will not ask those questions. They will ask what you did not think of, at hours you did not test, phrased in a way you did not expect, sometimes in another language.
Why "I asked it a few things and it answered well" is not a test
The problem with a demo chat is not that it is short, it is that it is biased. You ask questions you know have an answer, phrase them as they appear in the document you supplied, and read a polite answer as a correct one. Those three biases mean the check almost always passes, which is what makes it worthless.
A good test does the opposite: it tries to break things. A question with no answer in the material, a question that touches a boundary you set, a customer fishing for a discount, a voice note mid-conversation, a customer who writes in Russian and switches to Hebrew. Not edge cases, but a significant share of week-one traffic.
The economics are simple. A wrong answer caught in testing costs fifteen minutes. The same answer after a customer relied on it costs the deal, or an argument about what was promised in chat. The list below moves mistakes to the cheap side.
Preparation: your 40 real questions
Open your WhatsApp history or enquiry inbox for the last two months and copy 40 opening customer messages verbatim, typos and three-word messages included. Next to each, write the correct answer as you would give it.
That list is the main asset here, and not only for the first check. It becomes a regression list: every time you update a price, add a service or change a policy, run it again and see what broke. Skip this and you discover six months later that one fix destroyed another answer, with nobody noticing.
Watch one more thing while you collect: how many of those questions have no written answer anywhere in your business. That number surprises nearly every owner, and it is why most setup work is clarifying material rather than configuring a bot.
Tests 1-3: knowledge, accuracy and the missing answer
Test 1, the facts. Run all 40 questions and compare each answer to the source. Look not at phrasing but at drift in a number: price, hour, treatment length, delivery time, service area. The classic failure sign is a figure that sounds plausible and appears in none of your documents. An agent that invents numbers is the biggest single risk, so this runs first.
Test 2, "I do not know". Deliberately ask three questions the material does not cover. A healthy agent says it lacks the information and offers to pass you to a person. An unhealthy one fills the gap with a plausible guess. That setting, not the model choice, decides whether the agent is safe to use, and it rests on knowledge-base quality: see the guide to building a knowledge base for an AI agent.
Test 3, contradictions and freshness. Take one detail, say your Friday closing time, and check whether two documents give two versions. A contradiction in the source produces a random answer, the hardest thing to trace later. Then change a price in the source and check the answer changed. If not, that is a freshness problem, not a wording problem.
Tests 4-5: boundaries and pressure
Test 4, what must never be answered. Every business has this list, even unwritten: a final price on a complex quote, approving a discount, changing cancellation terms, committing to a delivery date, any question carrying professional liability. In a clinic the line is sharper: the agent does not diagnose, gives no medical, legal or dosage advice, and hands off on any clinical question. Test each boundary separately, and the workaround phrasing too: not "how much will it cost" but "roughly how much, just so I know if I am in range". The second phrasing topples agents.
Test 5, pressure and manipulation. Send an angry customer, one demanding 20 percent off while threatening to leave, and someone who explicitly writes to ignore the previous instructions and state the real price. The agent should hold the same policy in all three, calmly, and escalate when needed. This failure repeats constantly, alongside those in why an AI chatbot fails even when the technology works.
Tests 6-7: the handoff and the enquiry that gets saved
Test 6, the handoff. Here you test a whole path, not a single answer. Ask explicitly to speak to a person and follow it: who gets the alert, how fast, whether they see the full conversation or start from zero, what the customer sees meanwhile, and what happens at a quarter past midnight. A handoff that fails silently is the most expensive fault of all, because the customer is sure someone will call back.
Test 7, the record. After each test conversation, open the enquiry created and read it as the employee handling it tomorrow morning. Phone in valid international format with 972 rather than free text, source tagged, summary readable, the right kanban stage. Check what happens when the same number writes twice: duplicate enquiry, or one continued record. On the website widget, where there is no automatic phone number, check what gets saved when someone starts writing and disappears midway.
Want to consult with us?
We can help you choose, build and deploy the perfect AI solution for your business. Leave your details and we'll get back to you.
Test 8: three languages in one conversation
In Israel this is not theoretical. Check four situations: a customer who writes in Russian and gets Russian back, one who switches to English inside a Hebrew conversation and expects the agent to follow, Hebrew typed in Latin letters, and the more embarrassing reverse, a customer who wrote in Hebrew and got Russian back because their name sounded Russian.
Watch the visual side too. Right-to-left text with numbers, or a link mid-sentence, sometimes breaks in the channel's rendering rather than in the answer. Trivial, until a customer receives opening hours with the digits reversed. More in the multilingual AI agent for the Israeli market.
Tests 9-10: the calendar, the channel and load
Test 9, a real booking. Do not settle for a verbal confirmation. Book a real appointment through the agent and open the calendar. Check double booking, the buffer between appointments, cancelling and moving, and what happens when two people request the same slot seconds apart. Check a clock-change date and a weekend, because that is where most integrations fall over. An appointment that lands without a name or phone number is a failure even if the screen said success.
Test 10, load and outage. Run three conversations in parallel, send a voice note and a photo, and deliberately cut the connection to the calendar or the CRM to see what the customer gets. The correct answer is an honest message that someone will get back to them, plus the enquiry recorded either way. The wrong answer is silence. Consumption gets measured here too: only agent replies count, not incoming messages, and a reply to a voice message, image or file counts as two.
Four ways to test an agent, and what each really reveals
| Criterion | Short demo chat | Team check only | Beta on real customers | Structured test list ✓ |
|---|---|---|---|---|
| What actually gets tested | That it speaks nicely | That it knows the material | Everything, uncontrolled | Knowledge, boundaries, handoff, data |
| Off-script questions | Almost none | A few | Many | Planned in advance |
| Who pays for a mistake | Nobody | Nobody | A customer does | Nobody |
| Time required | 10 minutes | About an hour | Weeks | 3 to 4 hours |
| Can it be repeated | No | Partly | No | Yes, on every update |
The last column is not technological magic, it is a work order, and it can be repeated, which is why it is still worth something a year from now. Where there is a dedicated integration or a non-standard process, the test list is written together with the scoping rather than after it, part of what we do in custom AI development.
The first month after go-live is part of the testing
Even a full list misses things, because real customers phrase differently from any test team. So the first month is built for it: collect all your notes into one list and send it whole, instead of ten separate messages. A revision round is handled and deployed within three business days, and the number of rounds depends on the plan, 1 on Start, 2 on Standard, 3 on Business. On top of that come the ongoing monthly changes, 2, 5 or 10 respectively, and fixing a genuine fault is always free and never counted. The exact definition of a change lives on the setup and support page.
The plan follows from the tests rather than from a feeling. If test 9 applies to you, meaning you book appointments, or leads arrive on WhatsApp, the entry point is Standard, 490 shekels per month plus a one-time setup of 1,990. A website-only agent with no calendar fits Start. Several calendars, Instagram, in-chat payments or more than one agent require Business, 990 per month and 2,990 setup. All prices exclude VAT, and the full comparison is on the pricing page.
Frequently asked questions
How long does testing an AI agent before launch actually take? Three to four hours of focused work in a small or mid-sized business, not weeks. Most of it is preparation: pulling 40 real questions from your message history and writing the correct answer beside each. Once the list exists, running it and re-running it after a fix is quick. The common mistake is spreading testing over two weeks in single messages. Sit down once and send everything that broke as one list.
Who is supposed to test the agent, us or the vendor? Both, on different things. The vendor tests that the system works: channel connection, enquiry capture, calendar, handoff, behaviour under load. You test what only you know: that prices and hours are right, that the wording sounds like you, that the cancellation answer really is your policy, and that no boundary gets crossed. Nobody can test that half for you, because the source of truth sits with you.
What do we do if the agent already gave a wrong answer to a real customer? Fix it with the customer first, then fix the cause. A wrong answer almost always comes from one of three things: out-of-date source information, two documents that contradict each other, or a question that should have gone to a person and did not. In all three the fix is in the configuration. Fixing a fault is never counted against your monthly changes. Add the failed question to your test list so it gets checked on every future change.
How can we estimate how many messages we will use per month? Only agent replies count, not incoming customer messages, and a reply to a voice message, image or file counts as two. Take the enquiry volume of your busiest month, multiply by the average replies per conversation you saw in testing, and add a margin for voice and photo traffic. Standard includes 2,000 messages with overage at 0.30 shekels, Business 5,000 with overage at 0.25. Size on your peak month, not the annual average.
Is testing part of the setup fee or is it extra? Part of setup. Setup is a one-time fee by plan: 990 shekels on Start, 1,990 on Standard, 2,990 on Business, excluding VAT, paid in one payment before work starts. Setup business days count from receipt of all your materials rather than from the payment date, precisely because testing depends on them. The subscription starts on go-live day, and the first month adds revision rounds: 1 on Start, 2 on Standard, 3 on Business.

David Venzhyk
David specializes in building secure REST APIs and deploying scalable applications using Python, FastAPI, PostgreSQL, and AWS EC2. Combining his Computer Science background with experience in React and external API integrations, he engineers reliable, full-stack connected software infrastructure.