Playbook · 5 minute read
How to Build a Multilingual Chatbot
A multilingual chatbot detects the customer's language reliably, retrieves from content in that language or translates retrieved content with care, generates natively where the model is strong and translates where it is not, is evaluated separately per language with native reviewers, and escalates to human agents who speak that language. Quality varies by language and must be measured, not assumed.
The typical path to a multilingual chatbot is building one that works in English and switching on more languages, on the assumption that the model handles them. It does, unevenly, and the retrieval content, the evaluation sets, and the escalation paths were all built for English. The result is a bot that works well in one language and quietly poorly in the rest, with nobody measuring the difference. This guide covers building one that actually works across languages, drawing on FISTA Solutions' AI agents delivery for international customer operations. It complements how to build an ai chatbot and how to build an ai translation workflow.
Where does quality vary?
| Component | Varies by language because |
|---|---|
| Model generation | Training data volume differs enormously across languages |
| Language detection | Short messages and related languages confuse detectors |
| Retrieval | Content coverage differs; embeddings vary in cross-lingual quality |
| Register and tone | Formality conventions differ; a casual English tone may offend |
| Evaluation | Only measured where someone built a set |
| Escalation | Agent language coverage differs by shift and region |
Every row must be handled per language, and the first design decision is which languages the bot genuinely supports versus which it should decline gracefully.
How should language be detected?
Per message, with confidence, and with the ability to switch. Customers switch languages mid-conversation, code-switch within a message, and write short messages that detectors misclassify. Detection should run on each message, use the conversation history and the customer's known preference as priors, and express confidence so that a low-confidence detection on a short message defers to the prior rather than switching the bot's language on a single ambiguous word.
Closely related languages and shared scripts are the hard cases, and the evaluation set should include them.
Translate or generate natively?
Decide per language, on evidence. Native generation, where the model reasons and responds in the target language directly, produces more natural output and preserves nuance, and it works well where the model is strong in that language. Translation, where the model works in a pivot language and the input and output are translated, is more robust for languages where the model is weak, at the cost of naturalness and occasional translation errors.
Most deployments end up mixed: native generation for major languages where evaluation shows quality holds, and a pivot pattern for long-tail languages. The decision should be made per language from that language's evaluation set, and revisited as models change.
How does retrieval work across languages?
Best when content exists in each target language and is retrieved directly. Many organisations have help content only in their primary language, which leaves three options: translate the content corpus into each language and maintain it, which is the highest quality and highest cost; use multilingual embeddings to retrieve source-language passages for a target-language query and generate natively from them; or translate retrieved passages on the fly.
The second option works well for major languages and should be evaluated per language, because cross-lingual retrieval quality varies. Citations should point to the source document whatever the path, so the customer can see where the answer came from. See what is hybrid search.
What about register and cultural convention?
They matter more than translation accuracy for whether the bot feels acceptable. Formality expectations, honorifics, directness, and the handling of names differ across languages and markets, and a bot that is friendly in English can read as rude or absurd elsewhere. Native reviewers should define the register per language, and the evaluation rubric should score it.
How is it evaluated?
Per language, with native speakers, on real questions. The reference set for each language is built from actual customer messages in that language, not from the English set translated, because customers ask different things in different markets and phrase them differently. Native reviewers score correctness, groundedness, naturalness, and register. Scores are reported per language, never aggregated, because the aggregate hides the language that is failing. See the RAG evaluation methodology whitepaper.
How should escalation work?
To a person who speaks the language, with the conversation and its detected language passed through. Where the contact centre's language coverage varies by shift, the bot should know current availability and, where no agent for that language is available, say so and offer an alternative such as a callback, an email in that language, or a different channel. Handing a conversation in one language to an agent who speaks another is worse than an honest wait.
How is mixed-language input handled?
By expecting it. Customers write product names in English within another language, quote error messages verbatim, and switch languages when frustrated. The bot should respond in the customer's dominant language, preserve quoted identifiers and product names as written, and not treat a single foreign word as a language switch. The evaluation set should include code-switched messages, because they are common in most multilingual markets.
What does the build sequence look like?
One week deciding supported languages and assembling per-language reference sets from real messages with native reviewers. One week on detection with confidence and switching logic. Two weeks on retrieval per language, testing native versus cross-lingual per language. Two weeks on generation with per-language register defined and evaluated. One week on escalation routing by language and availability. Launch one language at a time, each on its own evidence.
What goes wrong?
Switching on languages without per-language evaluation. English reference sets translated and called multilingual. Retrieval that finds nothing because the content is English-only. A single register applied everywhere. Detection per session, so a customer who switches is answered in the wrong language. And escalation to a queue that does not speak the customer's language.
How FISTA Solutions helps
FISTA Solutions builds multilingual chatbots with per-message detection, retrieval and generation strategies chosen per language on evidence, register defined by native reviewers, evaluation sets built from real messages in each language, and escalation routed by language and availability, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies across 12+ countries with 99.9% uptime.
To serve customers in their language rather than in translated English, message FISTA on WhatsApp, or read how to build an ai chatbot.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why do multilingual chatbots fail after working in English?
Because model quality, retrieval coverage, and evaluation were all built for English and assumed to transfer. Models are weaker in many languages, knowledge content often exists only in English, and nobody evaluated the other languages with native speakers, so failures go unnoticed until customers complain.
02Should the chatbot translate or generate natively?
Generate natively where the model's quality in that language is strong, measured on your own evaluation set, and translate where it is not. Many deployments use native generation for major languages and a translate-generate-translate pattern for long-tail languages, with the choice made per language on evidence.
03How should retrieval work across languages?
Best with content in each target language, retrieved directly. Where content exists only in one language, multilingual embeddings can retrieve across languages, and the retrieved passage is translated or used as grounding for native generation, with the citation pointing to the source-language document.
04How is quality evaluated per language?
With a reference set per language built from real customer questions in that language, not translated from English, scored by native speakers on correctness, groundedness, naturalness, and register. Aggregate scores across languages hide the language where the bot is failing.
05How should escalation work?
Route to a human agent who speaks the customer's language, with the conversation and its detected language passed through, and where no such agent is available, say so honestly and offer an alternative channel rather than handing a Portuguese conversation to an English-only queue.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.