FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Glossary ¡ 5 minute read

What Is Chunk Overlap? Document Splitting for RAG Explained

Chunk overlap repeats a portion of text between adjacent chunks so that a passage split across a boundary remains retrievable intact. It costs index size and produces near-duplicate results, and it is a mitigation for arbitrary splitting rather than a substitute for splitting along the document's own structure.

By FISTA Solutions¡ AI-Native Engineering Team¡
What Is Chunk Overlap? Document Splitting for RAG Explained article cover

Chunking is the least discussed and most consequential decision in a retrieval system. Teams that adopt a default splitter with default parameters frequently spend weeks afterwards tuning prompts and embedding models to compensate for context destroyed at ingestion. This explainer covers what overlap does, what it costs, and what usually works better. It complements what is semantic chunking and how to improve rag accuracy, and reflects FISTA Solutions' approach in AI agents delivery.

Why does splitting lose context?

Because a fixed-length cut lands wherever the count runs out. A definition may sit on one side and its exception on the other. A table header may be separated from its rows. A clause may be split from the condition that qualifies it.

Retrieval then returns a chunk that is locally coherent and materially incomplete, and neither the retrieval system nor the model has any signal that something is missing.

ApproachBoundary qualityIndex sizeEffort
Fixed size, no overlapPoorSmallestMinimal
Fixed size with overlapBetterLargerMinimal
Paragraph-alignedGoodNormalLow
Structure-awareVery goodNormalModerate
Semantic splittingVery goodNormalHigher
Parent-child retrievalExcellentLargerModerate

How does overlap help?

By repeating a portion of each chunk in its neighbour, so a passage spanning a boundary appears complete in at least one chunk. A definition split in half by a cut will be whole in the overlapping chunk that straddles it.

It is a genuine improvement and a blunt one: it addresses the symptom of arbitrary boundaries rather than the cause.

What does it cost?

Index size, proportionally. Twenty percent overlap means roughly twenty percent more chunks, with corresponding storage, embedding, and query cost.

It also produces near-duplicate retrieval results. A query matching the overlapped region returns two chunks containing much of the same text, consuming context budget twice for one piece of information. Deduplicating overlapping results before assembling context is a small change with a real effect.

What works better?

Splitting where the document already divides. Headings, sections, paragraphs, list items, and table rows are boundaries the author placed, and they usually mark where one idea ends.

A chunk corresponding to a logical unit rarely needs overlap, because there is no arbitrary cut to repair. Most enterprise documents — policies, manuals, contracts, documentation — have clear structure that a splitter can follow if asked to.

How should chunk size be chosen?

From the documents and from measurement, not from a default. Reference material with short self-contained entries wants small chunks. Narrative or argumentative text wants larger ones, because the meaning spans paragraphs.

Chunk size also interacts with the embedding model's effective input length, which is often shorter than its stated maximum. Very long chunks embed imprecisely, because a single vector must represent too much.

What metadata should each chunk carry?

Source document, section heading, position, effective date, and entitlement scope at minimum. A chunk arriving at a model without its heading is frequently ambiguous — "this does not apply in the cases described above" is unusable without knowing what section it came from.

Metadata also enables filtering before ranking, which improves precision and is where entitlement must be enforced. See what is metadata filtering in rag.

What is parent-child retrieval?

Indexing small chunks for precise matching but returning their larger parent section to the model. Retrieval benefits from specificity while generation benefits from context, which resolves the tension that overlap only papers over.

It costs more storage and some complexity, and for documents where context matters it substantially outperforms tuning chunk size and overlap.

How should overlap be chosen?

By measuring retrieval recall on a labelled query set at several settings. The right value differs by corpus and by document type, and the measurement takes hours rather than days.

Choosing it by intuition, or by copying a value from a tutorial, is how systems end up with an index forty percent larger than necessary and no better at retrieval.

What should you do first?

Look at your documents. If they have headings and sections, split along them and measure before adding any overlap at all. Most corpora that appear to need aggressive overlap simply had their structure ignored at ingestion, and fixing that is cheaper and better than compensating downstream.

How does chunking affect cost?

Directly. More chunks means more embeddings at ingestion, more vectors in the index, and more candidates to rank at query time. A corpus chunked at half the size with twenty percent overlap holds well over twice the vectors of one chunked along its structure, which shows up in storage, query latency, and re-embedding cost when models change.

Structure-aware splitting is therefore cheaper as well as better, which is unusual enough to be worth stating plainly.

How FISTA Solutions helps

FISTA Solutions splits along document structure rather than character counts, attaches interpretable metadata to every chunk, uses parent-child retrieval where context matters, deduplicates overlapping results before assembling context, and chooses overlap from measured recall, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To fix retrieval at the ingestion step rather than the prompt, message FISTA on WhatsApp, or read how to improve rag accuracy.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why does splitting lose context?

Because a fixed-length cut falls wherever it falls, frequently mid-sentence or mid-argument. A definition on one side of the boundary and its qualifying condition on the other means retrieval returns half the answer, and the model has no way to know the rest existed.

02How much overlap is typical?

Commonly ten to twenty percent of chunk size, though the right figure depends on the corpus and should come from measurement. More overlap improves boundary recall and inflates index size and duplicate hits, so it is a trade rather than a setting with a best value.

03What does overlap cost?

Index size grows proportionally, with storage, embedding, and query cost following it, and retrieval returns near-duplicate chunks that consume context budget twice for one piece of information. Deduplicating overlapping results before assembling the context is a small change with a real effect.

04What is better than more overlap?

Splitting along the document's own structure: sections, headings, paragraphs, list items, and table rows. A chunk corresponding to a logical unit rarely needs overlap, because the boundary was placed where the meaning already ended rather than where a character count ran out.

05What metadata should chunks carry?

Enough to be interpretable alone: source document, section heading, page or position, effective date, and entitlement scope. A retrieved chunk reaching a model without its heading is frequently ambiguous in ways the model cannot detect.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project