Glossary · 5 minute read
What Is a Feature Store? Consistent ML Features Explained
A feature store manages the computed inputs to machine learning models: defining features once, serving them consistently to training and inference, and preventing the training-serving skew that arises when the same feature is computed differently in two places. It also enforces point-in-time correctness during training.
Feature stores solve a real and specific problem that is hard to see until it has already cost you a model's accuracy. They are also adopted routinely by teams whose scale does not justify them, adding infrastructure to solve a problem shared library code would have handled. This explainer covers what the problem is and how to tell which situation you are in. It complements what is ai data readiness and how to build a data quality agent, and reflects FISTA Solutions' approach in AI enablement delivery.
What problem does it solve?
Feature inconsistency. A model is trained on features computed by a batch pipeline in one language and served features computed by application code in another. The definitions drift — a different aggregation window, different handling of nulls, a different rounding rule — and the model receives inputs that differ subtly from what it learned on.
The result is a model that performed well in evaluation and performs worse in production, with no error anywhere to explain it.
| Problem | Without a feature store | With one |
|---|---|---|
| Feature definition | Duplicated in two places | Defined once |
| Training-serving consistency | Hope | Guaranteed by construction |
| Point-in-time correctness | Manual and error-prone | Enforced |
| Cross-team reuse | Reimplementation | Shared catalogue |
| Feature lineage | Usually absent | Recorded |
| Serving latency | Ad hoc | Purpose-built online store |
Why is skew so damaging?
Because it is silent. Nothing fails. The pipeline runs, the model serves, the metrics look healthy, and accuracy is quietly lower than evaluation promised. Diagnosing it requires comparing feature distributions between training and serving, which teams do only after suspecting the problem.
Defining features once and serving both paths from that definition removes the category rather than managing it.
What is point-in-time correctness?
Ensuring that a training example uses only the feature values that existed when that example occurred. It sounds obvious and is easy to get wrong: joining current customer attributes onto historical transactions gives the model information from after the event it is predicting.
The symptom is excellent offline performance that production cannot reproduce, because in production the future is genuinely unavailable. Feature stores enforce this at the join, which is where teams most often make the mistake by hand.
What are online and offline stores?
Two serving paths from one set of definitions. The offline store holds historical values for training, optimised for large batch reads across long time ranges. The online store holds current values for inference, optimised for single-key lookups in milliseconds.
Keeping them in sync from shared definitions is the core function. Their different performance characteristics are why a single database usually cannot serve both well.
Does this apply to generative AI systems?
Partly. Generative systems use fewer engineered numeric features, so the classic case applies less directly. But routing decisions, retrieval filters, personalisation signals, and eligibility flags are features in the same sense, and wherever the same signal is computed in two places the consistency argument holds.
The point-in-time discipline also applies to evaluating agents against historical scenarios, where it is equally easy to leak information that was not available at the time.
Does every team need one?
No, and this is worth saying plainly. A team with three models, one pipeline, and no cross-team sharing gets the consistency benefit from a shared library imported by both paths, at a fraction of the operational cost.
Feature stores earn their complexity when many models across multiple teams share features, when online serving latency requirements are strict, and when point-in-time joins are being done by hand and going wrong.
What are the costs?
Another system to operate, with its own availability requirements in the inference path. A learning curve. Migration effort for existing pipelines. And the ordinary tendency of any catalogue to accumulate stale entries, which means feature ownership and deprecation need to be managed like any other shared asset.
What should you do first?
Check whether you have skew. Compare the distribution of each feature as computed in training against the same feature as served in production, on the same population. If they match, the problem you would be buying a feature store to solve does not exist yet, and shared definitions in code will keep it that way for some time.
How does it relate to data quality?
Closely. A feature store guarantees that training and serving use the same definition; it does not guarantee that the definition is computing something correct from data that is current. A feature consistently computed from a table that stopped updating three days ago is consistently wrong.
Feature freshness monitoring therefore belongs alongside the store rather than being assumed from it, and it is the check most often missing. See how to build a data quality agent.
What about feature ownership?
Every feature needs an owner, because shared features become load-bearing for models their author never knew about. A change to an aggregation window made for one model's benefit can silently degrade three others, and without ownership and downstream visibility nobody discovers that until accuracy has already fallen.
How FISTA Solutions helps
FISTA Solutions diagnoses training-serving skew before recommending infrastructure, enforces feature definitions in shared code where scale does not warrant a store, implements point-in-time correct joins, and builds feature platforms where multi-team reuse and latency requirements justify them, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To fix feature consistency without over-building, message FISTA on WhatsApp, or read ai data readiness.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is training-serving skew?
When a feature is computed one way for training and another way at inference — different code, different window, different null handling. The model performs well in evaluation and worse in production, and the cause is invisible because both pipelines look correct.
02What is point-in-time correctness?
Ensuring that training examples use only feature values that were available at the time the example occurred. Without it, training accidentally uses future information, producing excellent offline results and a model that cannot reproduce them in production.
03What are online and offline stores?
The offline store holds historical values for training, optimised for large batch reads. The online store holds current values for inference, optimised for low-latency lookup. Both derive from the same definitions, which is what guarantees consistency.
04Does this apply to LLM systems?
Partly. Generative systems use fewer engineered numeric features, but retrieval filters, routing decisions, and personalisation signals are features in the same sense, and the consistency argument applies wherever the same signal is computed in two places.
05Does every team need one?
No. A team with a few models, one pipeline, and no cross-team reuse gets consistency more cheaply from shared library code. Feature stores earn their complexity at the scale where many models and teams share features.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.