Checklist ¡ 5 minute read
AI Monitoring and Alert Checklist: Catching What Matters
Availability monitoring stays green while an AI system produces wrong answers. Monitor quality signals, retrieval health, cost rate, latency at the tail, and agent behaviour, with alerts that reach someone who can act and thresholds tuned so they are not routinely ignored.
Availability monitoring stays green while an AI system produces wrong answers. This checklist covers the signals that actually indicate a problem, drawn from FISTA Solutions' AI agents operational work.
What should be monitored?
Six signal groups, only one of which conventional monitoring covers.
| Signal group | What it detects |
|---|---|
| Availability and errors | Infrastructure failure |
| Quality proxies | Output degradation |
| Retrieval health | Context failure |
| Cost rate | Runaway spend |
| Latency distribution | Experience degradation |
| Agent behaviour | Trajectory problems |
Quality signals
The signals that detect the failure conventional monitoring misses. See how to monitor AI quality in production.
- Human correction rate tracked where review exists
- Escalation rate tracked and trended
- User retry or rephrase rate tracked
- Validation failure rate on structured output tracked
- Refusal and abstention rate tracked
- Production sampling reviewed by a person on a schedule
- Quality metrics visible on the same dashboard as availability
Retrieval health
Degrades independently of the model. See RAG quality checklist.
- Retrieval hit rate measured on evaluation traffic
- Proportion of queries returning nothing relevant tracked
- Index lag behind source documents monitored
- Retrieval latency measured at the tail
- Corpus size and age distribution tracked over time
- Embedding model version recorded and monitored for change
- Alert when index refresh fails or falls behind
Cost monitoring
Rate-based, so problems are caught while small. See LLM cost control checklist.
- Spend rate tracked hourly
- Cost attributed per feature and workflow
- Anomaly detection on token volume
- Alert thresholds set well below the monthly ceiling
- Cost per task trended so gradual drift is visible
- Retry and loop counts monitored
- Cancelled or abandoned request cost measured
Latency and throughput
Measure what users experience.
- Time to first token tracked for streaming interfaces
- Total latency tracked at the ninety-ninth percentile
- Provider latency separated from your own processing time
- Rate limit responses counted and alerted
- Queue depth monitored where work is queued
- Timeout rate tracked
- Latency compared against the agreed budget per use case
Agent behaviour
Trajectories, not just outcomes. See agent trace analysis pipeline.
- Steps per task tracked and alerted on outliers
- Tool call distribution monitored for unexpected shifts
- Tool failure rate tracked per tool
- Tasks hitting the step limit counted
- Permission denials logged and reviewed
- Approval requests and their outcomes tracked
- Unusual action sequences surfaced for review
Alert hygiene
An alert nobody acts on is noise with a cost.
- Every alert routes to a person or rota that can act
- Every alert has a runbook entry
- Thresholds tuned so routine variation does not fire
- Alert volume reviewed regularly and noisy alerts fixed or removed
- Severity levels distinguish urgent from informational
- Alerts tested by triggering them deliberately
- Dashboards exist for investigation, not only alerts for detection
What are the most common failures?
Monitoring availability only. Cost alerts on monthly totals. Retrieval treated as part of model quality. Alerts routed to a channel nobody watches. And thresholds so sensitive that everything is filtered.
Who should own this?
The team operating the system owns its alerts and their tuning. Alerts configured by a platform team and routed to a product team that did not choose them are the ones that get ignored.
How often should it run?
Reviewed monthly for noise and coverage, and after every incident to ask what signal would have caught it earlier. New alerts should arrive from incidents rather than from a template.
What evidence should it produce?
Alert response records, time to detection per incident, and a trend showing alert volume falling as noise is removed. Those show monitoring is being maintained rather than accumulated.
What if you have no ground truth for quality?
Use proxies and sampling. Correction rate, escalation rate, and validation failures are all available immediately and all correlate with quality.
Combine them with periodic human review of a random sample. That gives a slower but genuine measure to calibrate the proxies against, which is enough to detect a real degradation. See how to build an agent evaluation harness.
What should you do first?
Add one quality proxy â correction rate or escalation rate â to the dashboard your team already watches. That single signal detects the failure availability monitoring cannot.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: quality proxies and retrieval health monitored alongside availability, with alerts routed to people who can act and thresholds tuned against real variation, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read how to monitor AI quality in production.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why is availability monitoring insufficient?
Because an AI system can respond promptly with wrong answers indefinitely. Every conventional signal stays green while the thing the system exists to do stops working.
02What are proxy quality signals?
Measurable behaviours correlated with quality: human correction rate, escalation rate, user retry rate, and the proportion of outputs failing validation. All are available in real time.
03Why alert on cost rate?
Because a runaway loop or a misconfigured retry can consume a month's budget in hours. Rate alerts catch it while it is small; monthly totals report it after the fact.
04What retrieval signals matter?
Hit rate, the proportion of queries returning nothing relevant, index lag behind the source, and retrieval latency. Each degrades independently of the model.
05What makes an alert useful?
It reaches someone with the context and authority to act, at a threshold tuned so it fires for real problems. Alerts that fire routinely get filtered and then the real one is missed.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.