Checklist · 4 minute read
AI Localization Checklist: Shipping in More Than One Language
AI systems that work in English frequently degrade in other languages without anyone noticing, because nobody evaluates in those languages. Localisation needs per-language retrieval, per-language evaluation with native reviewers, managed terminology, and correct locale formatting for dates, numbers, and currency.
AI systems that work in English frequently degrade elsewhere, and the degradation is invisible without per-language measurement. This checklist covers it, drawn from FISTA Solutions' AI enablement delivery work.
What has to be localised?
Six layers, not just the interface strings.
| Layer | What localisation requires |
|---|---|
| Interface text | Translation and review |
| Corpus and retrieval | Content in the query language |
| Prompts and instructions | Language-appropriate phrasing |
| Evaluation | Native-speaker cases per language |
| Formatting | Locale conventions |
| Support and escalation | People who speak it |
Language coverage
Decide explicitly which languages you support to what standard.
- Supported languages listed with a quality commitment each
- Languages where the system should decline identified
- Language detection tested including mixed-language input
- Behaviour defined for unsupported languages
- Regional variants distinguished where they matter
- Right-to-left languages handled if in scope
- Character encoding verified end to end
Corpus and retrieval
The most common cause of poor non-English answers. See RAG quality checklist.
- Corpus coverage per language measured
- Retrieval tested with queries in each supported language
- Cross-language retrieval behaviour decided deliberately
- Embedding model's language coverage verified
- Language recorded as metadata on every document
- Region-specific content distinguished from translated content
- Gaps per language identified and prioritised
Evaluation per language
The step that makes degradation visible. See how to build an agent evaluation harness.
- Evaluation cases created in each supported language
- Cases written natively, not translated from English
- Native-speaker domain experts score the results
- Quality reported per language, not aggregated
- Regression detection runs per language
- Edge cases specific to each language included
- Results compared across languages to find weak coverage
Terminology and voice
Managed centrally, not re-decided per piece.
- Glossary maintained per language
- Product and feature names decided per language
- Terms that should not be translated listed
- Regulated and legal phrasing confirmed per market
- Formality register decided per language and applied
- Glossary supplied to the system, not only to reviewers
- Terminology consistency checked automatically where possible
Formatting and locale
Immediately visible when wrong, and occasionally dangerous.
- Date format correct per locale and unambiguous
- Number and decimal separators correct
- Currency displayed with the right symbol and placement
- Address and postal formats correct
- Name order handled per convention
- Time zones handled explicitly
- Units of measure appropriate to the market
Support and operations
Escalation needs language coverage or it is not an escalation path.
- Escalation staffed in each supported language
- Support hours cover the relevant time zones
- Reviewers available per language for human-in-the-loop steps
- Incident communication prepared per language
- Feedback from users in each language reaches the team
- Per-language quality metrics on the operational dashboard
- Local regulatory requirements checked per market
What are the most common failures?
Evaluating only in English. Translating evaluation cases rather than writing them natively. A corpus in one language serving queries in several. Terminology decided per piece. And escalation with no coverage in the user's language.
Who should own this?
A named owner per language or market, with the overall system owner accountable for the coverage decision. Languages without an owner degrade first.
How often should it run?
Per-language evaluation on every release, full review quarterly, and a market review whenever a new language is added or a market's regulation changes.
What evidence should it produce?
Per-language evaluation results over time, glossary version history, and native-speaker review records. Aggregate quality figures hide exactly what this checklist exists to find.
What if you cannot evaluate in every language?
Support fewer languages properly rather than many badly. A system that declines a language politely is better than one that answers it incorrectly with confidence.
Where a language must be supported without full evaluation, say so internally, monitor user feedback closely in that language, and route more of it to human review. See why human oversight is a design problem.
What should you do first?
Run your evaluation suite in your second-largest language with a native speaker scoring it. The gap against English is usually larger than expected.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: quality measured and reported per language with native-speaker scoring, and terminology managed centrally rather than re-decided per piece, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI content review checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why does quality degrade in other languages?
Because models perform unevenly across languages and nobody measures it. The English evaluation passes, the system ships, and quality in other languages is unknown until users complain.
02What is the retrieval problem?
A corpus in one language queried in another may retrieve nothing relevant. Either the corpus needs localising, or retrieval needs to bridge languages deliberately rather than by accident.
03Who should evaluate?
Native speakers with domain knowledge, per language. Machine-assessed translation quality does not capture whether an answer is correct and appropriate in that market.
04Why manage terminology?
Because product names, technical terms, and regulated phrases must be consistent. Translating them afresh each time produces a system that names your own product three different ways.
05What about formatting?
Dates, numbers, currency, addresses, and name order differ by locale and errors are immediately visible. They also cause real mistakes when a date is read the wrong way round.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.