Glossary · 4 minute read
What Is Context Engineering? Managing What the Model Sees
Context engineering is the practice of deciding what information reaches a model on each call, in what order, and within what token budget. It covers retrieval selection, memory, tool outputs, and instructions. In production systems it affects accuracy, latency, and cost far more than prompt wording does.
Teams spend considerable effort refining prompt wording and comparatively little deciding what information the model actually receives. In production systems the second determines behaviour far more than the first. Context engineering is the discipline of making those decisions deliberately. This explainer covers what it involves and how to do it well, drawing on FISTA Solutions' AI agents delivery work. It complements the context engineering whitepaper and how to improve rag accuracy.
What does context engineering cover?
Everything in the model's input other than the instruction itself: retrieved documents and how many, conversation history and how much, tool outputs and their formatting, examples, schemas, and the order all of it appears in.
Each of these is a decision with consequences for accuracy, latency, and cost. Left undecided, they default to whatever the framework does, which is usually to include everything available.
| Context element | Typical default | Better practice |
|---|---|---|
| Retrieved chunks | Fixed top-k | Relevance-thresholded |
| Conversation history | Everything | Summarised beyond a window |
| Tool outputs | Raw | Filtered to what is needed |
| Instructions | Long and layered | Concise, at the edges |
| Examples | Many | Few, chosen for the case |
| Schemas | Full | Only the relevant parts |
Why does position matter?
Because attention across a long input is uneven. Models generally weight material near the beginning and end more heavily than material in the middle, an effect observed consistently enough to design around.
The practical rule is to place critical instructions and the most relevant evidence at the edges of the context, and lower-priority material in between. A crucial constraint buried halfway through a long context is measurably less likely to be followed.
Why is more context often worse?
Because irrelevant material competes with relevant material. Retrieving twenty chunks when five are relevant does not give the model more to work with; it gives it fifteen opportunities to ground an answer in something that merely resembles the question.
Long contexts also cost more and respond more slowly. The instinct to use the full window because it is available is one of the most reliable ways to build a system that is simultaneously expensive, slow, and less accurate than a tighter one.
What is a context budget?
An explicit allocation: this many tokens for system instructions, this many for retrieved evidence, this many for history, this much reserved for the response. Enforced, with a defined strategy for what gives when the budget is exceeded.
Without one, the failure is silent. Conversation history and tool outputs grow across a session until something is truncated — often the system instructions, because they are at the start — and behaviour changes without any error. See what is a context budget.
How should history be handled?
By summarising rather than accumulating. Keeping the last several turns verbatim and a running summary of what came before preserves continuity at a fraction of the token cost, and it preserves it more reliably, because a summary states conclusions that raw history only implies.
What must never be summarised away is anything the user will assume is remembered: decisions made, constraints stated, and corrections given.
What about tool outputs?
They are the most common source of silent budget consumption. A database query returning two hundred rows, an API response with deeply nested metadata, a file listing — each can consume more context than the entire rest of the input.
Tool outputs should be filtered and formatted for the model's actual need before insertion, which is application code rather than a model decision.
How is context engineering measured?
By task success on a fixed evaluation set, with ablation. Remove one context component, re-run the evaluation, and observe the effect. Components that make no measurable difference are cost and latency with no benefit, and they are common.
Cost per successful task is the complementary metric: a context change that improves accuracy by two points while tripling cost may or may not be worth it, and the decision needs both numbers.
How does this relate to agents?
Agents make it harder. Multi-step execution accumulates tool outputs, intermediate reasoning, and history across many calls, and naive implementations carry all of it forward. Deciding what each step actually needs — which is rarely everything from prior steps — is the difference between an agent that completes long tasks and one that degrades partway through.
What goes wrong most often?
Fixed top-k retrieval regardless of relevance. Unbounded conversation history. Raw tool output pasted into context. Critical instructions placed in the middle of long inputs. And no measurement, so nobody knows which parts of a carefully assembled context are contributing anything at all.
How FISTA Solutions helps
FISTA Solutions designs context budgets explicitly, thresholds retrieval by relevance rather than fixed counts, summarises history rather than accumulating it, filters tool outputs in application code, and measures context decisions by ablation against task success, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To make your context work harder and cost less, message FISTA on WhatsApp, or read the context engineering whitepaper.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How does this differ from prompt engineering?
Prompt engineering concerns the instruction's wording. Context engineering concerns everything else in the window: retrieved documents, conversation history, tool results, and their arrangement. In production systems the second determines outcomes far more than the first.
02Why does position matter?
Because models attend unevenly across long inputs, typically weighting the beginning and end more than the middle. Critical instructions or evidence buried in the centre of a long context are demonstrably less likely to influence the output than the same text placed at either edge.
03Is more context always better?
No. Adding marginally relevant material dilutes the signal, raises cost, increases latency, and can actively mislead. A system retrieving twenty chunks where five are relevant usually performs worse than one retrieving the five, not better.
04What is a context budget?
An explicit allocation of the available window between system instructions, retrieved evidence, conversation history, tool outputs, and the reserved response. Without one, history and tool results expand until something important is silently truncated mid-session, and behaviour changes with no error to show for it.
05How do you know the context is working?
Measure task success against a fixed evaluation set while varying what is included. Ablation — removing one component and re-measuring — is the most direct way to learn whether a given piece of context is contributing anything at all.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.