Playbook · 5 minute read
How to Build a Security Alert Triage Agent for SOC Teams
A security alert triage agent enriches alerts with asset, identity, and threat context, correlates related alerts into single incidents, proposes a disposition with calibrated confidence, and hands an assembled investigation to an analyst. Containment actions that affect production or user access remain under human authority.
Security operations centres are defined by a volume problem. Analysts receive more alerts than can be investigated, and most of the time spent on each is not investigation but enrichment — looking up who owns the asset, what the user normally does, whether the indicator is known, and what else fired nearby. Automating that layer is the highest-value intervention available, and it does not require the agent to decide anything. This guide covers building one, drawing on FISTA Solutions' AI agents work in security operations. It complements the AI security operations whitepaper and ai incident response. This article is general guidance, not legal advice.
Why start with enrichment?
Because it is the bulk of the work and none of the judgement. For each alert an analyst establishes asset ownership and criticality, the user's normal behaviour, recent related activity, threat intelligence on the indicators, and whether anything similar has been seen. That is several minutes per alert of lookups across five consoles.
Automating it changes the analyst's starting position from a bare alert to a contextualised one, which shortens every subsequent step. It is also low-risk: enrichment adds information and decides nothing.
| Triage stage | Automatable | Notes |
|---|---|---|
| Enrichment | Fully | Lookups, no judgement |
| Correlation | Largely | Rules plus similarity |
| Disposition proposal | With calibration | Confidence must be measured |
| Investigation assembly | Fully | Timeline and evidence |
| Judgement | No | Analyst |
| Containment | Narrow set only | Explicit agreement |
What does correlation actually change?
It converts alert volume into incident volume. A single lateral movement sequence generates alerts across endpoint, identity, and network telemetry, and analysts triaging them individually reach partial conclusions and close them separately. Correlated, it is one incident with a clear narrative.
Correlation should combine deterministic rules — same host, same user, same time window, same campaign indicators — with similarity over alert content. Pure similarity groups things that merely look alike; pure rules miss what the rules did not anticipate.
How should disposition confidence work?
Calibrated against outcomes. A system saying it is ninety percent confident an alert is benign should be right about ninety percent of the time, and that must be measured against analyst decisions and incident outcomes rather than asserted by the model.
Uncalibrated confidence is actively harmful. Analysts who learn the numbers mean nothing discount all of them, including the correct high-confidence signals, and the automation's value collapses. Calibration should be reported and tracked as a first-class metric.
Why hand off an investigation rather than a verdict?
Because analysts must verify, and a verdict without evidence makes them redo the work. The handoff should contain a timeline, the entities involved, the evidence gathered, the hypotheses considered, and explicitly what was checked and ruled out.
That last element is the one most often missing and the most valuable. Knowing that the agent checked the parent process, the user's travel records, and the asset's patch state — and what it found — saves the analyst from repeating those checks.
What containment can be automated?
A narrow set, agreed explicitly and in advance: isolating a non-critical endpoint on high-confidence malware detection, blocking a hash confirmed malicious by multiple sources, revoking a session on high-confidence credential compromise.
Anything affecting production services, broad user populations, or business-critical assets needs human authority. The reason is asymmetry: a false positive that isolates a production database causes an outage, and the agent cannot weigh that against the security benefit. See human in the loop ai explained.
How are false negatives handled?
As the metric that matters most and is hardest to see. An agent that confidently dispositions alerts as benign will eventually be wrong about a real intrusion, and nobody will notice at the time.
The controls are sampling — analysts review a random selection of auto-closed alerts — and retrospective analysis after every confirmed incident to check whether related alerts were auto-dispositioned. Without both, false negative rate is unmeasured and the system's real performance is unknown.
What about threat intelligence quality?
Variable, and it propagates. Low-quality feeds produce false positives that the agent will faithfully amplify at machine speed. Intelligence sources should be scored on their own precision, and enrichment should present source and confidence rather than treating all indicators as equal.
How does it integrate?
With the SIEM or detection platform as the alert source, the case management system for incidents, and the enrichment sources — identity, asset inventory, EDR, threat intelligence — through APIs. Analysts should work in their existing console. A separate triage interface fragments the workflow and loses adoption.
How is it evaluated?
On analyst minutes per incident, alert-to-incident ratio after correlation, disposition accuracy against analyst review, calibration error, false negative rate from sampling, and time to containment on true positives. Alerts processed is the metric that looks best and means least.
What does the build sequence look like?
Two weeks on enrichment integrations, which delivers value immediately and safely. Two weeks on correlation. Three weeks on disposition with a calibration harness built alongside, not afterwards. Two weeks on investigation assembly. Automated containment last, narrow, and only once false negative measurement is running.
What goes wrong?
Starting with disposition because it is the impressive part. Uncalibrated confidence. Verdicts without evidence. No false negative sampling. Automated containment before the evidence supports it. Amplified low-quality intelligence. And measuring alert throughput.
What does it cost to run?
Enrichment is cheap and high volume. Disposition reasoning is the expensive component, and routing only ambiguous alerts to larger models while handling clear cases with small ones keeps costs proportionate. Budget also for the calibration and sampling work, which is analyst time and is not optional.
What should you do first?
Time your analysts. Measure how long enrichment takes per alert across a week and multiply by volume. In most SOCs that single number justifies the enrichment layer immediately, and it gives the programme a baseline to be measured against later.
How FISTA Solutions helps
FISTA Solutions builds SOC triage systems with comprehensive enrichment, rule-and-similarity correlation, calibrated disposition confidence, investigation handoff including what was ruled out, false negative sampling, and tightly bounded automated containment, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To give analysts their time back without losing detections, message FISTA on WhatsApp, or read the AI security operations whitepaper.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why is enrichment the right starting point?
Because it is the largest, most repetitive, and least judgemental part of triage. Looking up asset ownership, user context, recent activity, threat intelligence, and related alerts takes minutes per alert and requires no analyst expertise, yet it consumes most of a shift.
02What does correlation change?
It converts alert volume into incident volume, which is the number that actually matters. Twenty alerts from one lateral movement sequence are one incident, and analysts triaging them individually will reach twenty partial conclusions instead of one correct one.
03How should confidence be expressed?
Calibrated against outcomes, not asserted. A disposition claiming ninety percent confidence should be right about ninety percent of the time, measured. Uncalibrated confidence is worse than none, because analysts learn to distrust every number the system produces.
04Why hand off an investigation rather than a verdict?
Because analysts must be able to verify, and a verdict without evidence forces them to redo the work. An assembled investigation — timeline, entities, evidence, what was checked and ruled out — makes review fast and keeps judgement human.
05What containment can be automated?
A narrow, explicitly agreed set: isolating a non-critical endpoint, blocking a confirmed-malicious hash, revoking a session on a high-confidence credential compromise. Anything affecting production services or broad user access needs human authority. This is general guidance, not legal advice.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.