FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Checklist · 4 minute read

AI Localization Checklist: Shipping in More Than One Language

AI systems that work in English frequently degrade in other languages without anyone noticing, because nobody evaluates in those languages. Localisation needs per-language retrieval, per-language evaluation with native reviewers, managed terminology, and correct locale formatting for dates, numbers, and currency.

By FISTA Solutions· AI-Native Engineering Team·
AI Localization Checklist: Shipping in More Than One Language article cover

AI systems that work in English frequently degrade elsewhere, and the degradation is invisible without per-language measurement. This checklist covers it, drawn from FISTA Solutions' AI enablement delivery work.

What has to be localised?

Six layers, not just the interface strings.

LayerWhat localisation requires
Interface textTranslation and review
Corpus and retrievalContent in the query language
Prompts and instructionsLanguage-appropriate phrasing
EvaluationNative-speaker cases per language
FormattingLocale conventions
Support and escalationPeople who speak it

Language coverage

Decide explicitly which languages you support to what standard.

  • Supported languages listed with a quality commitment each
  • Languages where the system should decline identified
  • Language detection tested including mixed-language input
  • Behaviour defined for unsupported languages
  • Regional variants distinguished where they matter
  • Right-to-left languages handled if in scope
  • Character encoding verified end to end

Corpus and retrieval

The most common cause of poor non-English answers. See RAG quality checklist.

  • Corpus coverage per language measured
  • Retrieval tested with queries in each supported language
  • Cross-language retrieval behaviour decided deliberately
  • Embedding model's language coverage verified
  • Language recorded as metadata on every document
  • Region-specific content distinguished from translated content
  • Gaps per language identified and prioritised

Evaluation per language

The step that makes degradation visible. See how to build an agent evaluation harness.

  • Evaluation cases created in each supported language
  • Cases written natively, not translated from English
  • Native-speaker domain experts score the results
  • Quality reported per language, not aggregated
  • Regression detection runs per language
  • Edge cases specific to each language included
  • Results compared across languages to find weak coverage

Terminology and voice

Managed centrally, not re-decided per piece.

  • Glossary maintained per language
  • Product and feature names decided per language
  • Terms that should not be translated listed
  • Regulated and legal phrasing confirmed per market
  • Formality register decided per language and applied
  • Glossary supplied to the system, not only to reviewers
  • Terminology consistency checked automatically where possible

Formatting and locale

Immediately visible when wrong, and occasionally dangerous.

  • Date format correct per locale and unambiguous
  • Number and decimal separators correct
  • Currency displayed with the right symbol and placement
  • Address and postal formats correct
  • Name order handled per convention
  • Time zones handled explicitly
  • Units of measure appropriate to the market

Support and operations

Escalation needs language coverage or it is not an escalation path.

  • Escalation staffed in each supported language
  • Support hours cover the relevant time zones
  • Reviewers available per language for human-in-the-loop steps
  • Incident communication prepared per language
  • Feedback from users in each language reaches the team
  • Per-language quality metrics on the operational dashboard
  • Local regulatory requirements checked per market

What are the most common failures?

Evaluating only in English. Translating evaluation cases rather than writing them natively. A corpus in one language serving queries in several. Terminology decided per piece. And escalation with no coverage in the user's language.

Who should own this?

A named owner per language or market, with the overall system owner accountable for the coverage decision. Languages without an owner degrade first.

How often should it run?

Per-language evaluation on every release, full review quarterly, and a market review whenever a new language is added or a market's regulation changes.

What evidence should it produce?

Per-language evaluation results over time, glossary version history, and native-speaker review records. Aggregate quality figures hide exactly what this checklist exists to find.

What if you cannot evaluate in every language?

Support fewer languages properly rather than many badly. A system that declines a language politely is better than one that answers it incorrectly with confidence.

Where a language must be supported without full evaluation, say so internally, monitor user feedback closely in that language, and route more of it to human review. See why human oversight is a design problem.

What should you do first?

Run your evaluation suite in your second-largest language with a native speaker scoring it. The gap against English is usually larger than expected.

How FISTA Solutions helps

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: quality measured and reported per language with native-speaker scoring, and terminology managed centrally rather than re-decided per piece, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI content review checklist.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why does quality degrade in other languages?

Because models perform unevenly across languages and nobody measures it. The English evaluation passes, the system ships, and quality in other languages is unknown until users complain.

02What is the retrieval problem?

A corpus in one language queried in another may retrieve nothing relevant. Either the corpus needs localising, or retrieval needs to bridge languages deliberately rather than by accident.

03Who should evaluate?

Native speakers with domain knowledge, per language. Machine-assessed translation quality does not capture whether an answer is correct and appropriate in that market.

04Why manage terminology?

Because product names, technical terms, and regulated phrases must be consistent. Translating them afresh each time produces a system that names your own product three different ways.

05What about formatting?

Dates, numbers, currency, addresses, and name order differ by locale and errors are immediately visible. They also cause real mistakes when a date is read the wrong way round.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project