Web & Mobile ┬╖ 5 minute read
Caching Strategy: Where to Cache and What to Invalidate
Caching is easy to add and hard to invalidate. The decisions that matter are which layer to cache at, what the cache key includes, how invalidation happens, and what the system does when the cache is cold or stale.
Caching is straightforward to add and difficult to invalidate correctly, which is why cached systems serve wrong data more often than slow ones do. This guide covers the decisions that matter, drawing on FISTA Solutions' web and mobile work.
What layers are available?
Several, each faster than the last and less able to reflect a change.
| Layer | Suits |
|---|---|
| Browser cache | Static assets, immutable content |
| CDN | Public content, edge-personalisable |
| Application cache | Computed results, session-scoped data |
| Query result cache | Expensive repeated queries |
| Database buffer | Automatic, not your concern |
| Semantic cache | Repeated AI queries |
How do you choose the layer?
The closest to the user that can still serve a correct response.
A static asset belongs in the browser and the content delivery network. A personalised page belongs in an application cache or is assembled at the edge from cached fragments. Data that changes per request should not be cached at all.
Caching too far from the user wastes the opportunity; caching too close serves stale or wrong content. The correctness constraint decides it.
What makes a good cache key?
One that includes every dimension the response varies on and nothing else.
Missing a dimension serves the wrong content to someone тАФ the classic case being a cached page that varies by user but keys only on the URL. Including too many produces a key space so large that nothing ever hits.
List the dimensions explicitly: path, query, locale, device class, authentication state, experiment assignment. Then remove the ones that do not actually change the response.
Which invalidation strategy?
Time-based expiry where staleness is tolerable, event-driven invalidation where correctness matters, and stale-while-revalidate where you want both speed and freshness.
Stale-while-revalidate serves the cached copy immediately and refreshes in the background, which suits most content well: users get a fast response and the next one is current.
Event-driven invalidation is stronger and harder. It requires knowing every cache entry affected by a change, which is straightforward for keyed entries and difficult for derived or aggregated ones.
How do you handle cache stampedes?
With request coalescing, jittered expiry, and background refresh.
When a popular entry expires, every concurrent request misses and hits the origin simultaneously. That can take down the origin, which is a cache causing an outage rather than preventing one.
Coalescing means one request fetches and the rest wait. Jitter spreads expiry so entries do not all lapse together. Background refresh replaces entries before they expire. Use all three for anything popular.
What should happen with a cold cache?
The system should work, more slowly, without falling over.
After a deploy, a flush, or an incident, the cache is empty and every request reaches the origin. A system sized on the assumption of a warm cache will fail at exactly that moment.
Size the origin for a cold cache at realistic traffic, or warm the cache deliberately before taking traffic. Discovering this during an incident is common and avoidable.
What about caching authenticated content?
Possible and requires care. The cache key must include enough of the authentication context that one user cannot receive another's content.
The safest pattern is caching the shared parts and assembling the personalised parts per request, or caching per user with a short lifetime. Caching a full authenticated page keyed only on the path is the failure that makes the news.
Verify with a test that attempts to retrieve one user's cached content as another. See multi-tenant SaaS architecture.
What are the common mistakes?
Cache keys missing a dimension. Too many dimensions, so nothing hits. Manual purging as the strategy. No stampede protection. Origin sized for a warm cache. And caching authenticated content on the path alone.
How do you test it?
Test cache hit rates against realistic traffic, test invalidation end to end, test cold start under load, and test that authenticated content cannot cross users.
The last is the one worth automating as a permanent check, because it is the failure with the largest consequence.
What does it cost to operate?
Cache infrastructure is cheap relative to the origin capacity it saves. The real cost is engineering time spent on invalidation and on debugging stale data.
That cost rises with the number of layers. Each additional cache is another place a change must propagate, which is an argument for fewer layers with clearer rules.
What should you measure?
Hit rate per layer and per key pattern, origin load, stale-content incidents, and latency at the tail with a cold cache.
What about caching AI responses?
Exact-match caching is safe and helps where the same question arrives repeatedly, which in support and internal knowledge systems is often.
Semantic caching covers more and needs care: two similar questions can have different correct answers, particularly where the underlying data changes. Set similarity thresholds conservatively and invalidate on corpus updates. See what is a semantic cache.
When is this the wrong approach?
When the data changes on every request or when correctness cannot tolerate any staleness. A cache that must be invalidated on every write is overhead with no benefit, and the honest answer is to make the origin faster.
What should you do first?
List every cache layer between your users and your data, and for each one say how it is invalidated. Any layer without a clear answer is where your stale data is coming from.
How FISTA Solutions helps
FISTA Solutions builds and operates production systems through web and mobile, AI enablement, and staff augmentation: cache layers chosen for correctness as well as speed, invalidation designed explicitly rather than relying on expiry and manual purges, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To scope this work, message FISTA on WhatsApp, or read CDN strategy guide.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Where should you cache?
At the layer closest to the user that can serve a correct response: the browser, a content delivery network, an application cache, or the database query layer. Each is faster and less able to reflect recent changes.
02What makes a good cache key?
One that includes everything the response varies on and nothing else. Too few dimensions serves wrong content to someone; too many produces a cache that never hits, which is complexity with no benefit.
03What invalidation strategies work?
Time-based expiry for content that tolerates staleness, event-driven invalidation where correctness matters, and stale-while-revalidate where both matter. Manual purging is a fallback rather than a strategy.
04What is a cache stampede?
Many requests missing the cache simultaneously and all hitting the origin, which can take it down. It happens on expiry of a popular entry and on cold start after a deploy, and it needs explicit handling.
05What is the worst failure mode?
Serving stale data confidently. An expired price, a revoked permission, or a superseded policy served from cache is worse than a slow response, because nothing indicates it is wrong.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.