FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Glossary ┬╖ 5 minute read

What Is Tokenization? How Models Read Text, Explained

Tokenization is the process of splitting text into the units a model processes, which are usually sub-word fragments rather than whole words. Token counts determine context limits and billing, vary substantially between languages and content types, and differ between model families, which makes counting them per model necessary rather than optional.

By FISTA Solutions┬╖ AI-Native Engineering Team┬╖
What Is Tokenization? How Models Read Text, Explained article cover

Tokenization is the least glamorous part of working with language models and one of the most frequent sources of production surprises: a cost forecast that is wrong by a factor of three for a non-English market, a truncation that breaks structured output, a context limit hit far earlier than expected. This explainer covers what tokenization is and where it bites. It complements what is a context budget and llm cost per task benchmarking, and reflects how FISTA Solutions approaches cost control in AI enablement work.

What is a token?

A fragment of text from a fixed vocabulary the model was trained with. Common words are usually a single token; rarer words split into several pieces; unusual strings split into many. Punctuation and whitespace are tokens too, and leading spaces are typically attached to the following word.

The model never sees characters or words directly. It sees a sequence of token identifiers, which is why some tasks that look trivial to a human are structurally awkward for a model.

Content typeApproximate efficiencyNote
English proseMost efficientTokenizers optimised for it
Other Latin-script languagesModerately efficientVaries by language
Non-Latin scriptsOften several times costlierDirect cost impact
Source codeInefficientPunctuation and identifiers
JSON and XMLVery inefficientStructural characters
Numbers and IDsInefficientSplit into fragments

Why sub-words rather than words?

Because a word-level vocabulary cannot cover an open-ended language. Names, typos, technical terms, and new coinages would all be unknown. Sub-word tokenization lets any string be represented from a bounded vocabulary by decomposing unfamiliar sequences into familiar pieces.

The trade-off is that the model's view of text is fragmented in ways that do not align with human intuition about words.

Why does language affect cost so much?

Because tokenizer vocabularies are learned from training corpora that over-represent certain languages, usually English. A language with less representation gets fewer dedicated tokens, so its text decomposes into more fragments.

The practical effect is significant: the same message can cost several times more in one language than another, and a product serving multiple markets will see costs that do not track user counts. Any multilingual cost model must account for this explicitly rather than applying an English-derived average.

Does tokenization affect what models can do?

Yes, in specific ways. Character-level tasks тАФ counting letters in a word, reversing a string, certain arithmetic manipulations тАФ are structurally harder because the model sees fragments. Numbers in particular tokenize inconsistently, which contributes to arithmetic unreliability.

This is not a gap that more training closes cleanly; it follows from the representation. The engineering answer is to handle such tasks in code rather than asking the model to perform them.

Can token counts be estimated?

Roughly, for English prose, using a words-to-tokens ratio. The estimate fails badly for code, structured data, non-Latin scripts, and heavily punctuated text тАФ precisely the content where budgets matter most.

Any system enforcing context limits or cost controls should count with the actual tokenizer for the model in use. Tokenizers differ between model families, so a count from one provider's tokenizer is not valid for another's.

What bugs does tokenization cause?

Truncation is the main one. Cutting a context at a token boundary can split a JSON structure, a code block, or a sentence, and the model receives something malformed without any error being raised. Truncation logic should be aware of structure, not only of token counts.

Budget miscalculation is the second: a system sized on English assumptions runs out of context far earlier on other input. And streaming logic that assumes token boundaries correspond to word boundaries produces odd partial output in the interface.

How does this affect chunking?

Chunk sizes are usually specified in tokens, and the same character count produces different chunk counts across languages. A retrieval system tuned on English documents will behave differently on a multilingual corpus, with more chunks per document and different boundary placement. Testing chunking per language is worth the effort where the corpus is mixed.

How should teams handle it?

Count with the right tokenizer, budget per language rather than on an average, truncate with structure awareness, and keep character-level tasks out of the model. None of this is difficult; all of it is commonly skipped, and each omission surfaces as an unexplained cost or a rare malformed output. See how to build an ai cost dashboard.

How do tokenizers differ between providers?

Each model family ships its own tokenizer with its own vocabulary, learned from its own training corpus. The same paragraph can produce noticeably different counts across families, which matters when comparing prices: a lower per-token price with a less efficient tokenizer may cost more per unit of work.

Cost comparisons should therefore be made per task on representative content, not per token on a rate card. That single discipline changes model selection decisions more often than any benchmark score does.

What about output tokens?

They are typically billed at a higher rate than input and are the component teams forget to control. A prompt that invites a long answer costs more on every call for the life of the system, and asking explicitly for brevity or constraining with a structured schema is one of the cheapest optimisations available.

How FISTA Solutions helps

FISTA Solutions builds cost models that account for per-language tokenization, enforces context budgets with real tokenizers, implements structure-aware truncation, and keeps character-level operations in deterministic code, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies across 12+ countries.

To forecast AI costs accurately across languages, message FISTA on WhatsApp, or read how to build an ai cost dashboard.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why not just split on words?

Because a fixed word vocabulary cannot cover every word in every language, including names, typos, and technical terms. Sub-word tokenization represents any string from a limited vocabulary by breaking unfamiliar words into familiar fragments, which is what makes models work across open-ended text.

02Why do some languages cost more?

Because tokenizers are trained on corpora that over-represent some languages. Text in an under-represented language or script may use several times as many tokens to express the same meaning, which directly multiplies both cost and effective context consumption.

03Does tokenization affect quality?

Indirectly and sometimes substantially. Tasks requiring character-level reasoning тАФ counting letters, reversing strings, some arithmetic тАФ are harder because the model sees fragments rather than characters. It is a structural limitation, not a training gap.

04Can I estimate tokens as words divided by a constant?

Only roughly, and only for English prose. Code, structured data, non-Latin scripts, and heavily punctuated text all break the heuristic badly. Any system with cost controls or context limits should count with the actual tokenizer for the model in use.

05What bugs does tokenization cause?

Truncation that cuts mid-structure, producing malformed JSON or half a sentence; budget calculations that undercount for non-English input; and streaming logic that assumes token boundaries align with characters or words. All are quiet failures rather than errors.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.

Start a project