FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook · 6 minute read

How to Build a Twilio AI Agent

A Twilio AI agent receives inbound messages and calls through webhooks, maintains conversation state externally keyed by channel and sender, responds within each channel's constraints, and escalates to people with context. Consent, sender registration, and messaging regulation govern what may be sent to whom, and voice adds a latency budget that shapes the whole design.

By FISTA Solutions· AI-Native Engineering Team·
How to Build a Twilio AI Agent article cover

Twilio sits beneath a large share of the world's customer messaging and calling, which makes it the natural substrate for conversational agents: inbound SMS, WhatsApp, and voice all arrive as webhooks, and outbound goes through one API. It is also where messaging regulation is enforced most directly, through sender registration, consent requirements, template approval, and carrier filtering that will silently drop non-compliant traffic. This guide covers building agents that work within that, drawing on FISTA Solutions' AI agents delivery on messaging and voice. It complements how to build a whatsapp ai agent and how to build an ai voice agent for call centers. This article is general guidance, not legal advice.

How do the channels differ?

ChannelConstraintsAgent implications
SMSShort, plain text, carrier filtering, sender registrationConcise responses; compliance before capability
WhatsAppRich content, template approval, 24-hour service windowBusiness-initiated messages need approved templates
VoiceReal-time, latency-sensitive, interruptionStreaming speech pipeline, sub-second turns
RCS and othersVaries by marketTreat as extensions of the SMS model

Treating these as one channel with different transports produces a poor agent on each. SMS responses that run to three messages annoy; WhatsApp business-initiated messages outside the service window without a template fail; voice responses that take three seconds lose the caller.

How is conversation state handled?

Externally. Each Twilio webhook is a stateless HTTP request carrying the message and sender; the agent must look up where this conversation stands. The state store is keyed by channel and sender identifier and holds structured state, meaning the goal, the facts established, the step reached, and the outstanding question, rather than a raw transcript.

Two disciplines matter. Bound the history the agent reasons over, so a long-running WhatsApp thread does not grow its context unboundedly. And expire state deliberately, so a customer returning after a month starts fresh rather than being treated as mid-conversation. See the agent memory architecture whitepaper.

What compliance is required before anything works?

Consent, sender registration, and opt-out handling, and they are prerequisites rather than features. In several markets SMS traffic from businesses requires brand and campaign registration, and unregistered or non-compliant traffic is filtered by carriers without notification. WhatsApp requires business verification, approved message templates for anything the business initiates, and adherence to a policy that governs content.

Opt-out must work immediately and universally: a stop keyword must halt all messaging to that number across every workflow, and the agent must never message a number that has opted out. Consent records must exist for every recipient and must be checked before sending, not assumed. Message content is personal data and falls under data protection obligations.

Projects that build the agent first and address compliance at launch discover that launch is where compliance blocks them.

How does voice change the design?

Fundamentally, through latency. A voice agent must hear the caller stop, transcribe, retrieve, reason, synthesise, and begin speaking within roughly a second, and every component competes for that budget. Streaming speech recognition, streaming synthesis, small and fast models, pre-warmed retrieval, and interruption handling that stops the agent mid-sentence and discards the unspoken remainder are all required rather than optional.

Twilio's media streams provide the real-time audio path; the agent's pipeline must be built for it from the first line rather than adapted from a text agent. See the voice AI in the enterprise whitepaper.

What should the agent handle?

Bounded, high-volume interactions with clear resolution: appointment confirmation and rescheduling, order and delivery status, account balance and payment reminders with consent, verification codes and identity checks, and simple requests within defined policy. Each has a measurable resolution and a natural escalation point.

What it should not handle early: complaints, anything involving vulnerability, negotiations, and anything where a wrong answer has financial or safety consequences.

How does escalation work?

In the same channel, with context. A customer in a WhatsApp conversation should be handed to a person who continues the WhatsApp conversation with the history visible, not told to call a number. A caller on a voice line should be transferred with a whisper or screen pop that tells the person what was discussed.

Triggers should be generous: explicit request, repeated failure, frustration language, anything outside policy, and any distress signal. Twilio's routing capabilities support handoff to agent desktops; the engineering task is ensuring the state and history travel with the customer.

How is it evaluated?

Per channel, against real conversations. For messaging: resolution verified by absence of repeat contact, escalation rate and quality, and compliance incidents, which should be zero. For voice: transcript accuracy on domain terms, turn latency including the tail, interruption handling, task completion, and caller satisfaction measured separately from human-handled calls. See how to evaluate a voice agent.

What does the build sequence look like?

One to two weeks on compliance: consent model, sender registration, template approval, opt-out handling. One week on webhook infrastructure and external state. Two to three weeks on the first bounded workflow in one channel, usually messaging before voice. One week piloting with escalation staffed. Voice follows, with its own latency-driven build. Each additional channel reuses state and escalation patterns.

What goes wrong?

Compliance addressed at launch. State held in the webhook cycle or in raw transcripts. Voice built as a text agent with speech bolted on. Opt-out that works in one workflow and not another. Escalation that changes channel. Templates not approved, so business-initiated WhatsApp messages fail silently. And carrier filtering discovered through customers reporting they never received anything.

How should outbound campaigns be handled?

Separately from conversational agents and with more caution, because outbound is where regulation is strictest and where a mistake reaches thousands of people at once. Business-initiated messages need consent on record, an approved template on WhatsApp, appropriate sending hours, and rate control so a campaign does not trip carrier spam thresholds.

The agent's role in outbound is usually to handle the replies rather than to initiate: a reminder goes out as a template, and the agent handles the confirmation, reschedule, or question that comes back. That division keeps the initiation under human control and the conversation under the agent's, which is the right allocation.

How FISTA Solutions helps

FISTA Solutions builds Twilio agents with compliance settled first, external structured state, channel-specific design for SMS, WhatsApp, and voice, latency-engineered voice pipelines, and same-channel escalation carrying full context, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To put a conversational agent on messaging and voice that stays compliant, message FISTA on WhatsApp, or read how to build a whatsapp ai agent.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01How does a Twilio agent receive and respond to messages?

Through webhooks Twilio calls when a message or call arrives, to which the agent responds either synchronously with instructions or asynchronously by sending through the API. Conversation state is stored externally keyed by channel and sender, because each webhook is stateless.

02How do SMS, WhatsApp, and voice differ for an agent?

SMS is short, unformatted, and carrier-regulated with strict sender rules. WhatsApp supports rich content but requires approved templates for business-initiated messages and a 24-hour service window. Voice needs real-time speech handling with sub-second latency and interruption support.

03What compliance applies to messaging agents?

Consent to receive messages, opt-out handling that works immediately, sender registration such as brand and campaign registration for SMS in some markets, WhatsApp business policy including template approval, and data protection for message content. Confirm requirements with counsel.

04How should conversation state be managed?

In an external store keyed by channel and sender identifier, holding the structured state of the interaction rather than a raw transcript, with a bounded history window, expiry, and the ability to resume after a gap. Webhooks are stateless and state cannot live in them.

05How does escalation work across channels?

By handing the conversation to a person in the same channel, with the history and the agent's understanding of the request visible to them, so the customer continues the conversation rather than starting over on the phone or by email.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project