Before You Put an AI Agent in Front of Customers: The Checklist
Before an AI agent talks to your customers: where answers come from, what it does when it doesn't know, handoff, logging, testing, the off switch and privacy.
An AI agent can serve your customers without putting you at risk if you've first answered, in writing, the eight questions below. Each one you can't answer is a reason to wait another week, not to cancel. If the channel is WhatsApp, Meta's rules apply on top; we covered them in WhatsApp AI agents and won't repeat them here.
1. Where do the answers come from?
From your documents and your data, not from what the model "knows" about the world. Before launch, ask for the list of sources: the catalog, the policies, the FAQ, and the live data it connects to (inventory, calendar, order status). Whatever isn't on that list, the agent doesn't answer.
And ask for every answer to be traceable to its source. That's how we designed Tlamatini, a tax and legal advisor in our portfolio: every answer carries the Mexican statute or case law it comes from and passes an anti-hallucination check before it's shown. For a business the sources are simpler, but the principle is the same: if you can't see where an answer came from, you can't fix it.
2. What does it do when it doesn't know?
Says so. "I don't have that information. Want me to connect you with someone on the team?" is a good answer; a made-up price or an improvised delivery date is not. Language models tend to fill gaps with something that sounds right, which is why the rule has to be written and tested, not assumed.
Also define the list of topics the agent never answers even when it thinks it knows: complaints, billing, anything medical or legal. There it guides, takes the details and hands off.
Simple test: ask it something that isn't in its sources and watch what it does.
3. How does it hand off to a person?
Three things need to be defined:
- When. Whenever the customer asks, at any point; when the agent detects anger, urgency or a topic on the banned list; and when the customer repeats the same question because they're not getting an answer.
- To whom. A named person during business hours and a rule for after hours: what the customer is told, when they'll get a reply and who sends it.
- With what. Whoever picks up the conversation sees everything that was said and everything the agent promised. If the customer has to start over, the handoff didn't work.
4. What gets logged?
Every conversation in full: what the customer asked, what the agent answered, which sources it used, when it handed off and what happened next. Without that log you can't tell whether it works, or defend yourself when a customer says the agent promised them something.
Decide who has access to the log and how long it's kept. And measure a few things: how many conversations ended the way you wanted (a booking, an order, a resolved question), how many went to a person and which ones it couldn't answer. That last list is next week's work plan.
5. Did your team try to break it?
Before a customer sees it, your team messages it for a week trying to make it fail. The usual real questions, but also:
- Telling it to ignore its instructions or to "talk as if you were someone else." It's the risk that OWASP's list for applications built on language models puts first: prompt injection.
- Asking for another customer's details, or for it to repeat its internal instructions.
- Pushing, insistently, for a discount or a condition that doesn't exist.
- Writing with typos, in Spanish, with voice notes, with photos.
- Asking about a banned topic in a roundabout way.
Whatever fails gets fixed before launch, and every failure becomes a test that runs again whenever anything changes: the sources, the model or the instructions. If whoever is building it can't show you those tests, it isn't ready yet.
6. How do you switch it off?
With a switch anyone on your team can use, without calling the vendor and without waiting. When it's off, messages go to a person or get an honest note about when they'll be answered; they never go unanswered.
Put a spending cap on it too. If the AI model charges per use, a bug, a conversation stuck in a loop or someone with bad intentions can burn through a month's budget in a night; it's another risk on OWASP's list. A daily limit with an alert covers most of it. And rehearse the shutdown once before launch, like a fire drill.
7. What data does it collect, and where does it send it?
An agent asks for names, phone numbers, order details and sometimes an ID. Three things to check:
- Your privacy notice covers that use. If you serve customers in Canada, the federal law is PIPEDA, which applies to private-sector organizations in commercial activities and calls for, among other things, consent, safeguards and openness about what's done with the information. In Mexico, it's the federal data protection law for private parties (LFPDPPP). If you have customers in both countries, your notice has to cover both.
- Where the messages go. Which AI provider receives them, in which country they're stored, and whether its terms allow using them to train models. Ask for it in writing.
- What it shouldn't ask for. If it doesn't need the ID, it shouldn't ask. Less data, less risk.
8. Whose name are the accounts in?
The AI provider, the hosting, the WhatsApp number and account, the domain and the code repository: all in your company's name from day one, with your email and your card. If whoever built it disappears, the agent is still yours and another team can pick it up.
How to know it's ready
It's ready when you can show these eight things in writing: the list of sources, the list of banned topics, the handoff rule with a name and hours, the log, the week of testing with what failed and how it was fixed, the switch and the spending cap, the updated privacy notice, and the accounts in your name.
Then launch small: one topic, or after hours only, with a weekly review. If something goes wrong, you switch it off and normal service resumes.
If you'd like us to go through it with you
If you already have an agent, or are about to hire someone to build one, we'll go through it with you using this same list and tell you what's missing, even if someone else built it. See how we build AI agents or reach us through contact. And if you're not yet sure your case calls for an agent, start with when a WhatsApp agent makes sense or what to automate first.
- AI
- AI Agents
- Customer Service