FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Glossary · 5 minute read

What Is Prompt Compression? Reducing Context Cost Explained

Prompt compression shortens the context sent to a model while preserving the information needed for the task, through summarisation, selective inclusion, or token-level pruning. It reduces cost and latency and always loses something, which is why retrieving less rather than compressing more is usually the better fix.

By FISTA Solutions· AI-Native Engineering Team·
What Is Prompt Compression? Reducing Context Cost Explained article cover

Context costs money and latency, which makes compression attractive, and every compression technique loses information. The useful question is not how much can be removed but whether what was removed mattered — and frequently the better answer is to put less in rather than to compress what should not have been there. This explainer covers both. It complements what is a context budget and how to reduce ai costs, and reflects FISTA Solutions' approach in AI agents delivery.

What are the techniques?

Summarising conversation history. Filtering tool output to the fields the task needs. Selecting retrieved passages by relevance threshold rather than fixed count. Removing redundancy where several passages say the same thing. And token-level pruning, which drops tokens a model estimates carry little information.

They differ enormously in risk, and they are frequently discussed as if interchangeable.

TechniqueRiskTypical saving
History summarisationLowLarge in long sessions
Tool output filteringVery lowLarge where tools are verbose
Relevance-thresholded retrievalLowModerate, improves accuracy
Cross-passage deduplicationLowModerate
Extractive summarisation of passagesModerateModerate
Token-level pruningHighLarge, needs evaluation

Why is better retrieval preferable?

Because compressing bad retrieval preserves its badness in fewer tokens. Twenty passages of which five are relevant, compressed by half, is still mostly irrelevant content — now harder to read and equally misleading.

Retrieving the five relevant passages instead is shorter, cheaper, and more accurate. Compression is a reasonable response to context that genuinely needs to be there; it is the wrong response to context that should never have been retrieved.

What is the safest form?

History summarisation. Keeping the last several turns verbatim with a running summary of what came before preserves continuity at a small fraction of the tokens.

It frequently improves accuracy as well, because a summary states conclusions that raw history only implies — that the user is asking about a specific account, that a constraint was established, that an approach was rejected. Those are easier for the model to use than a transcript.

What about tool output?

Filtering it is the largest easy win in most systems. A database query returning two hundred rows, a deeply nested API response, or a directory listing can consume more context than the entire rest of the input.

Passing only the fields the task needs is application code, deterministic, lossless with respect to what matters, and frequently halves context on tool-heavy workloads. It is the first thing to do and it is regularly skipped.

How risky is token-level pruning?

More than its savings suggest. Removing tokens a model estimates as low-information can strip negations, qualifiers, and numbers — exactly the tokens that change meaning. The output remains fluent, which makes the damage hard to spot.

It can work with careful evaluation on a specific task. It should not be applied generally without measuring what it does to accuracy on that task.

How should it be measured?

By task success and cost per successful task. Compression ratio is the metric most likely to be reported and it describes effort rather than benefit. A system compressing sixty percent while failing more often has increased cost per outcome. See what is cost per task.

What should you do first?

Measure where your context actually goes, by component, on real traffic. In most systems one component dominates unexpectedly — usually tool output or unbounded history — and addressing that specific component achieves more than general compression would.

Does it affect accuracy?

Always, to some degree. The question is whether what was removed mattered for the task, which is measurable and rarely measured. Any compression step deployed without an accuracy comparison is an assumption presented as an optimisation.

Where does it belong in the pipeline?

After retrieval selection and before assembly, applied to components rather than to the finished prompt. Compressing an assembled context treats instructions, evidence, and history identically, when they have very different tolerance for loss.

Component-level compression lets history be summarised aggressively while instructions and the current question are left untouched, which is both safer and more effective.

What should be measured before and after?

Task success on a fixed evaluation set, cost per successful task, and latency. All three, because compression trades between them and improving one while degrading another is common. A compression change that reduces cost by a third and task success by five points is a decision, not an optimisation, and it should be made knowingly.

Who should own the compression logic?

The team that owns the system's quality, because compression is a quality decision wearing a cost decision's clothes. Placed with a cost-reduction initiative and measured on savings, it reliably compresses past the point where accuracy suffers.

What is the realistic order of work?

Filter tool output, threshold retrieval by relevance, summarise history, and only then consider lossy compression of what remains. Teams that start at the end of that list spend effort on the technique with the highest risk and the smallest remaining opportunity.

How FISTA Solutions helps

FISTA Solutions reduces context by improving retrieval and filtering tool output before applying compression, summarises history rather than accumulating it, evaluates any lossy compression against task success, and measures cost per successful task rather than compression ratio, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 47% efficiency gains.

To cut context cost without losing what matters, message FISTA on WhatsApp, or read what is a context budget.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What are the main techniques?

Summarising conversation history, filtering tool output to needed fields, selecting fewer retrieved passages by relevance threshold, removing redundancy across passages, and token-level pruning that drops low-information tokens. They differ greatly in risk.

02Why is better retrieval preferable?

Because compressing twenty mediocre passages into a shorter form preserves their mediocrity. Retrieving the five relevant ones instead is shorter, cheaper, and more accurate, and it addresses the cause rather than the symptom.

03What is the safest form?

History summarisation. Keeping recent turns verbatim and a running summary of earlier ones preserves continuity at a fraction of the tokens, and it often improves accuracy by stating conclusions that raw history only implied.

04What about tool output?

Filtering it is the largest easy win in most systems. A query returning two hundred rows or a deeply nested API response can dominate the context, and passing only the fields the task needs is application code rather than a model decision.

05How should compression be measured?

By task success and cost per successful task, not by compression ratio. A system compressing sixty percent while failing more often has not saved anything, and the ratio is the metric most likely to be reported.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project