FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Cost · 5 minute read

Anomaly Detection Cost: False Positives Are the Real Expense

Anomaly detection cost is dominated by investigating alerts rather than by building or running the detector. False positive rate determines that cost directly, and a system producing more alerts than the available investigation capacity has negative value because genuine anomalies are lost in the volume.

By FISTA Solutions· AI-Native Engineering Team·
Anomaly Detection Cost: False Positives Are the Real Expense article cover

Anomaly detection is inexpensive to build and expensive to operate, and the expense is entirely in human attention. Every alert consumes investigation time, which means the false positive rate rather than the detector's sophistication determines both the cost and whether the system gets used at all. This guide covers that economics, drawing on FISTA Solutions' AI enablement work. It complements how to build a data quality agent and how to build a security alert triage agent.

Why does investigation dominate?

Because every alert needs a person to determine whether it matters. The detector itself runs on a schedule at negligible cost; the recurring expense is the human time examining what it flagged.

That cost scales with alert volume rather than with data volume, which means a detector processing far more data but raising fewer alerts is cheaper to operate than a noisy one on a small dataset.

FactorEffect on costControllable
False positive rateDominantYes, through tuning
Baseline qualityLarge, via false positivesYes
Investigation capacitySets the ceilingOrganisationally
Alert enrichmentReduces time per alertYes
Detection computeSmallYes
False negative rateHidden riskOnly if measured

Why do false positives determine usability?

Because a system generating more alerts than can be investigated trains people to dismiss them. That dismissal becomes habitual, and once it does the genuine anomalies are dismissed alongside the noise.

At that point the system has negative value: it costs money, consumes attention, and provides false assurance that monitoring is in place. Alert fatigue is not a user experience problem; it is a failure of the control.

What makes baselines difficult?

Seasonality. Most business and system metrics vary by hour, day of week, and season, and a fixed threshold fires constantly at normal peaks while missing genuine anomalies during troughs.

Seasonality-aware baselines need weeks of history before they mean anything, which is a lead time worth stating to sponsors expecting immediate value. A detector deployed on Monday and alerting constantly by Wednesday will be muted by Friday.

How should alert volume be set?

To match investigation capacity. If two analysts can examine twenty alerts a day, a detector producing two hundred is misconfigured whatever its sensitivity.

Setting volume to capacity is a design constraint rather than a compromise, and it forces the useful question: which anomalies matter most. Prioritising by consequence rather than by statistical unusualness is what makes a capacity-constrained detector useful.

How can time per alert be reduced?

By enrichment. An alert arriving with the relevant context — what the metric normally does, what else moved at the same time, what changed recently, what happened last time this fired — takes a fraction of the time to assess.

That enrichment is engineering work done once and saves investigation time on every alert thereafter, which makes it the highest-return improvement available in most monitoring operations.

What about false negatives?

The hidden risk. A detector tuned to reduce alert volume may be missing genuine anomalies, and that is invisible without deliberate measurement.

Sampling what was not flagged, and retrospectively checking whether known incidents were detected, are the only ways to know. Neither is expensive and both are routinely omitted, which means most detectors' real performance is unknown.

What should you do first?

Count how many alerts your current monitoring raises and how many are investigated. If the second number is much smaller than the first, the system is already in the failure mode, and tuning to capacity will improve outcomes immediately without any new detection capability.

What determines whether an anomaly matters?

Consequence rather than statistical unusualness, and the two correlate weakly. A metric three standard deviations from its mean may be entirely benign; a small movement in a critical measure may be serious.

Encoding that judgement — which metrics matter, what magnitude of change is consequential for each, and what the downstream effect is — is domain work with the people who own the process. It is what turns a statistical detector into an operational control, and it is the part most often skipped because the statistics are easier to specify.

How does this scale across an estate?

Poorly, if each detector is configured independently. Organisations that deploy anomaly detection across many metrics without a shared approach to baselines, thresholds, enrichment, and routing end up with many noisy detectors and no coherent alerting.

Building the detection framework once — baselines, tuning, enrichment, routing, and false negative sampling — and applying it across metrics is substantially cheaper and produces alerting people can actually work with.

Who should own the tuning?

The team that investigates the alerts, because they bear the cost of noise and the risk of misses. Tuning owned by whoever built the detector optimises for detection rate; tuning owned by the investigating team optimises for usable alerting, which is the outcome that matters.

How FISTA Solutions helps

FISTA Solutions tunes anomaly detection to available investigation capacity, builds seasonality-aware baselines with honest lead times, enriches alerts to reduce investigation time, prioritises by consequence rather than statistical unusualness, and measures false negatives by sampling, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To build detection people actually act on, message FISTA on WhatsApp, or read how to build a data quality agent.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why does investigation dominate?

Because every alert requires a person to determine whether it matters. The detector runs cheaply; the human time spent examining what it flagged is the recurring cost, and it scales directly with alert volume rather than with data volume.

02Why do false positives determine usability?

Because a system generating more alerts than can be investigated trains people to dismiss them. Once that happens the genuine anomalies are dismissed too, and the system has negative value — it costs money and provides false assurance.

03What makes baselines difficult?

Seasonality. Most metrics vary by hour, day, and season, and a fixed threshold fires constantly at normal peaks while missing genuine anomalies in troughs. Seasonality-aware baselines need weeks of data before they mean anything.

04How should alert volume be set?

To match available investigation capacity. If two people can investigate twenty alerts a day, a detector producing two hundred is misconfigured regardless of its sensitivity. Tuning to capacity is a design constraint rather than a compromise.

05What about false negatives?

They are the hidden risk and are invisible without deliberate measurement. A detector tuned to reduce alert volume may be missing genuine anomalies, and sampling what was not flagged is the only way to know.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project