FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Trends ¡ 6 minute read

The Context Window Arms Race and Why It Misses the Point

Larger context windows are useful, but they do not remove the need to decide what a model should see. Cost scales with tokens sent, latency rises with input length, and accuracy degrades on material buried in the middle. Selection matters more as windows grow, not less.

By FISTA Solutions¡ AI-Native Engineering Team¡
The Context Window Arms Race and Why It Misses the Point article cover

Context windows have grown by orders of magnitude, and with each expansion teams conclude that retrieval engineering is now unnecessary. The economics and the accuracy data say otherwise. This piece explains why, drawing on FISTA Solutions' AI agents production work.

What actually changed?

Larger windows shift some constraints and leave others untouched.

What people assumeWhat holds in production
Retrieval becomes unnecessarySelection becomes more important
Just send the whole corpusCost scales per call, forever
More context means better answersAccuracy degrades in the middle
Latency is unaffectedLatency rises with input length
Simpler architectureSame architecture, worse economics
One prompt fits everythingTask-specific selection still wins

Why does cost dominate the argument?

Because input tokens are billed on every single call, not once.

A system that sends a hundred thousand tokens of context to answer each question pays for those tokens every time somebody asks. At a thousand questions a day, that is a hundred million input tokens daily for what may be a few hundred tokens of actually relevant material.

The engineering cost of building a selection layer is paid once. The token cost of not building one is paid continuously and grows with adoption, which is exactly the wrong shape for a system you want people to use more. See how to reduce AI costs.

What happens to accuracy?

It degrades unevenly, and the degradation is hardest to detect where it matters most.

Models use material at the start and end of a long input more reliably than material in the middle. A critical clause buried at the sixty percent mark of a long document is meaningfully less likely to influence the answer than the same clause at the top.

This is not a failure mode that announces itself. The model produces a confident, fluent answer that happens to have missed something, which is considerably worse than an obvious error. Evaluation is the only way to see it. See how to build an agent evaluation harness.

What about latency?

It rises with input length, which matters for anything interactive.

Processing a large input takes time before the first token of output appears. For a user waiting on a response, that shows up directly as the interface feeling slow, and it is the part of latency that streaming cannot hide.

For batch analysis nobody is waiting on, this is irrelevant. For a support interface or an agent in a workflow, it is the difference between usable and not. See how to optimize AI latency.

So when is long context the right tool?

When the volume is low and the material genuinely needs to be seen whole.

Analysing a single long contract, reviewing a codebase for a one-off question, or summarising a lengthy transcript are all cases where sending everything is both correct and cheaper than building infrastructure.

Prototypes are the other case. Sending everything is the fastest way to learn what the system actually needs, and the selection layer can be built once that is known. Treating it as a permanent architecture is where teams go wrong.

What does the selection discipline look like?

Deciding, per request, what the model sees — and measuring whether that decision was right.

That means retrieval, ranking, and assembly, with instrumentation showing whether the material needed to answer was actually present. The failure to diagnose is not a bad generation; it is a correct generation from an incomplete context.

Teams that measure retrieval quality separately from answer quality find and fix problems that teams measuring only the final output never locate. See how to improve RAG accuracy.

Where is this heading?

Windows will keep growing, and the economics will keep favouring selection.

Each expansion makes a wider class of one-off tasks practical without infrastructure, which is genuinely valuable. It does not change the arithmetic for high-volume production systems, where per-call cost is the constraint.

The likely equilibrium is that long context becomes the default for exploration and low-volume work, while production systems continue to select deliberately — not because they cannot send everything, but because sending everything is the expensive way to get a worse answer.

What is the counter-argument?

The strongest counter is that engineering time is more expensive than tokens for many teams, and that is often true at low volume. A team shipping an internal tool used forty times a day should send everything and move on. The argument holds where volume is high enough that per-call cost compounds, or where accuracy on buried material matters.

What does this change for engineering teams?

It means context engineering is a durable skill rather than a stopgap. Retrieval, ranking, chunking, and assembly remain the work, with the emphasis shifting from fitting within a limit to choosing well within a large budget.

It also means measurement of context quality — was the needed material present? — belongs in the evaluation suite alongside answer quality.

What does this change for buyers?

It means treating claims about context size as a feature rather than an answer. The question to ask a vendor is how they decide what the model sees and how they measure whether that decision was correct.

It also means modelling cost at your actual volume rather than at demonstration volume, because the two differ by orders of magnitude.

What should leaders do about it now?

Model your per-call token cost at realistic volume and compare it against the engineering cost of a selection layer. That single calculation usually settles the architectural argument.

Then require that retrieval quality be measured separately from output quality, so the team can see which half is failing.

Does this apply to agents?

More acutely. An agent accumulates context across many turns, so an approach that sends everything grows without bound as the conversation continues.

Agents need active context management — summarising earlier turns, dropping irrelevant tool output, and reassembling what matters for the current step. That is the same selection discipline applied over time rather than over a corpus. See AI agent production readiness checklist.

How will you know if this is happening?

Watch for per-call token cost rising faster than usage, for accuracy complaints that trace to material present but unused, and for latency creeping upward as prompts grow. All three indicate selection is being skipped.

How FISTA Solutions reads this

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: context selection measured separately from answer quality so teams can see which half is failing, and cost modelled at real volume before architecture is chosen, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To discuss what this means for your roadmap, message FISTA on WhatsApp, or read how to improve RAG accuracy.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Does a larger context window replace retrieval?

No. It changes what retrieval must do — from finding a few passages to selecting the right subset — but sending everything is expensive, slower, and less accurate than sending what matters.

02Why does accuracy degrade in long inputs?

Models attend unevenly across position. Material at the beginning and end is used more reliably than material in the middle, which means burying a critical fact in a long document reduces the chance it is used.

03What does long context actually cost?

Input tokens are billed on every call. A system sending a hundred thousand tokens per request pays that repeatedly, which at volume dwarfs the engineering cost of building proper retrieval.

04When is a long context genuinely right?

For one-off analysis of a large document, for low-volume workflows where engineering time is scarcer than tokens, and for prototypes where you are still learning what the system needs.

05What should teams build instead?

A selection layer that decides what enters the context for each request, with measurement of whether the right material was included. That discipline improves outcomes regardless of how large the window becomes.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project