Whitepaper · 9 minute read
Agent Reliability Engineering: Running AI Agents in Production
Agent reliability engineering is the practice of keeping AI agents dependable in production: defining service levels that include correctness, cataloguing failure modes that differ from ordinary software, running evaluation continuously rather than at release, instrumenting every step, bounding actions with guardrails and rollback, and operating a response model for failures that are silent rather than loud.
Software reliability engineering matured around a specific assumption: systems fail visibly. They return errors, exceed timeouts, exhaust memory, or stop responding, and the discipline built its alerting, dashboards, and on-call practice around detecting and responding to those signals. AI agents break that assumption. An agent can return a well-formed response in two hundred milliseconds with a two hundred status code and be completely wrong, and no conventional monitor will notice. This whitepaper defines the practices that close that gap. It draws on FISTA Solutions' production operations across AI agents deployments and complements the AI observability whitepaper and ai agent observability.
What is agent reliability engineering?
The discipline of keeping agents dependable in production. It borrows the structure of site reliability engineering, meaning service levels, error budgets, observability, incident response, and blameless review, and extends it to the failure modes that only probabilistic, tool-using systems exhibit. Its central claim is that correctness is an operational property, not a release-time property, because agent behaviour changes when models, prompts, tools, retrieval content, or user inputs change, which is continuously.
What is the failure taxonomy?
| Failure mode | What happens | Why conventional monitoring misses it |
|---|---|---|
| Plausible incorrectness | Fluent, confident, wrong answer | Returns 200 with normal latency |
| Groundedness loss | Answer not supported by retrieved sources | Retrieval succeeded; text looks cited |
| Retrieval degradation | Index stale or ranking degraded | Query returns results, just the wrong ones |
| Silent tool failure | Tool returns empty or partial data, agent proceeds | Tool call succeeded technically |
| Wrong action | Action within permissions but not appropriate | Policy allowed it |
| Loop and runaway | Agent repeats steps, burning cost | Only visible in cost, often late |
| Drift after provider change | Model updated, behaviour shifted | Nothing in your code changed |
| Prompt injection | Untrusted content redirects the agent | Input looked like ordinary data |
| Escalation failure | Agent should have handed off and did not | No error; the user simply gave up |
| Cost blowout | Context growth or retries multiply spend | Latency and errors normal |
The first row is the defining one. In ordinary software, output that looks right usually is right. In agents, looking right is what the model optimises for, which inverts the relationship between appearance and correctness.
What service levels should an agent carry?
Availability and latency remain, but they are the least informative. The set that matters adds:
- Task success rate: the proportion of interactions achieving the intended outcome, measured against an evaluation set and, where possible, verified downstream.
- Groundedness: for retrieval-based answers, the proportion of claims supported by cited sources.
- Action safety: policy violations or incorrect actions per thousand actions taken.
- Escalation rate and quality: too low means the agent is trapping users; too high means it is not resolving.
- Cost per task: with an alert threshold, because cost regressions are often the first visible sign of a behavioural regression.
Each gets a target and an error budget. When the budget is spent, changes stop and the team fixes quality before shipping features, which is the mechanism that keeps reliability from being traded away under delivery pressure. Service level design is in what is an slo for ai systems.
Why must evaluation run continuously?
Because the inputs change when nothing in your codebase does. Providers update models. Retrieval corpora are edited by people who do not know an agent depends on them. Upstream systems change their data formats. User behaviour shifts seasonally. A test suite that ran at release and passed says nothing about today.
Continuous evaluation means running a reference set on a schedule and on every change to any behaviour-affecting component, plus sampling live traffic for automated and human scoring. The reference set is a maintained asset with an owner, updated from production failures and retired when cases become obsolete. Practice is in what is continuous evaluation and what is a golden dataset.
What does observability require?
Traces at step granularity. For every interaction: the input, the resolved prompt, the retrieved context with document identifiers and scores, the model and exact version, each tool call with arguments and results, intermediate reasoning artifacts where available, the final output, the cost, and the latency of each step, all joined by a request identifier. Anything less makes post-hoc diagnosis guesswork, and agent diagnosis is almost entirely post-hoc because the failure was not visible when it happened.
On top of traces sit the aggregate signals: evaluation scores over time, groundedness rates, tool error and empty-result rates, escalation rates, cost distributions, and retrieval score distributions, which drift before quality does and therefore serve as a leading indicator. See what is distributed tracing for ai and how to build an agent trace analysis pipeline.
How do guardrails bound the damage?
By limiting what an agent can do, not by trying to make it always right. Permissions scoped to the agent's role. Action limits with value thresholds above which approval is required. Rate and volume caps that stop a loop before it becomes an incident. Output validation against schemas and business rules before anything is acted on. Confirmation steps for irreversible actions. And a kill switch that an on-call engineer can use without a deployment.
Blast radius is the design question: if this agent behaves as badly as it possibly could within its permissions, what is the worst outcome, and is that acceptable? Where the answer is no, the permissions are wrong. Guardrail design is in ai agent guardrails and permissions in how to design tool permissions for ai agents.
What does rollback cover?
Everything that changes behaviour, each versioned and revertible independently: prompts and system instructions, model selection and version pins, tool definitions and their schemas, retrieval indexes and their content snapshots, configuration and thresholds, and application code. A rollback capability that covers only code leaves the most common regression sources untouched, since most agent regressions come from prompt edits, provider model updates, or content changes.
Two practices make this real. Pin model versions rather than accepting provider defaults, so provider updates become a scheduled change with evaluation rather than a surprise. And snapshot retrieval indexes so a content change can be reverted as a unit. See what is a canary deployment and what is a shadow deployment.
How should changes be deployed?
Progressively, with evaluation gates. A behaviour-affecting change passes the reference evaluation, then runs in shadow against live traffic where the comparison is possible, then canaries to a small traffic share with quality metrics compared to the control, then expands. The gate is quality, not absence of errors.
For agents taking actions, shadow mode means executing the reasoning without the action and comparing proposed actions to those taken by the current system or by people. That catches action-selection regressions before they reach a customer account.
What does incident response look like?
Different from an outage. The detection is usually a metric drift, a sampled evaluation drop, a spike in escalations, or a customer report, not a page from a health check. The first response question is not "is it up" but "what changed", and the answer is usually a prompt edit, a provider model update, a content change, or an upstream data change.
Practical playbooks cover: quality regression, in which the team identifies the changed component, reverts it, and confirms via evaluation; wrong action taken, in which the team stops the agent, identifies affected records, remediates them, and only then investigates; prompt injection, in which the team isolates the content source and reviews what the agent accessed; and cost spike, in which the team caps and diagnoses whether it is volume or behaviour. Each ends with a blameless review that adds the case to the evaluation set, which is the mechanism that prevents recurrence. See ai incident response and the AI incident postmortem template.
What does on-call look like for agents?
The rotation needs people who can read a trace and judge whether an output was acceptable, which is a different skill from restarting a service. Runbooks specify how to disable an agent, how to revert each component, how to identify affected interactions, and who to notify when actions were taken in customer or financial systems.
Alerting should page on action-safety violations, cost anomalies, hard failures, and sustained evaluation score drops, and should ticket rather than page on softer quality signals, because the fatigue from paging on noisy quality metrics destroys the rotation faster than any other mistake. On-call design is in how to set up ai on call.
How does this scale across many agents?
Through shared infrastructure rather than per-agent heroics. A common gateway with version pinning and cost attribution, a shared tracing schema, a shared evaluation harness with per-agent reference sets, a shared guardrail library, and a single inventory recording every agent, its owner, its permissions, its service levels, and its current evaluation scores. Organisations running twenty agents without this end up with twenty different reliability practices and no ability to answer basic questions during an incident. See the agentic AI governance whitepaper.
What does maturity look like?
| Level | Detection | Evaluation | Rollback | Response |
|---|---|---|---|---|
| Initial | Customer reports | Manual spot checks | Code only | Ad hoc |
| Managed | Basic metrics, some alerts | Reference set at release | Code and prompts | Runbook exists |
| Defined | Traces plus quality metrics | Continuous plus sampling | All components versioned | Playbooks per failure type |
| Optimised | Leading indicators, drift alerts | Gated deployment, shadow, canary | Automated, tested regularly | Blameless review feeding evaluation |
Most organisations sit between initial and managed while believing they are at defined, because their agents have not yet failed in a way they noticed.
How FISTA Solutions delivers this
FISTA Solutions builds and operates production agents with service levels that include correctness, step-level tracing, continuous evaluation, versioned rollback across prompts, models, tools, and indexes, and incident playbooks written for silent failures, delivered through AI enablement, AI agents, and forward deployed engineers who set up the reliability practice inside client teams. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To run agents you can depend on, message FISTA on WhatsApp, or read the AI observability whitepaper.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is agent reliability engineering?
The operational discipline of keeping AI agents dependable in production, combining service level definition that includes correctness, failure mode analysis, continuous evaluation, step-level observability, guardrails and rollback, and incident response designed for silent quality failures rather than outages alone.
02How do agent failures differ from software failures?
Ordinary software fails loudly: errors, timeouts, crashes. Agents often succeed technically while being wrong, producing fluent, plausible output or taking a defensible-looking wrong action. Nothing alerts, dashboards stay green, and the failure surfaces through customers or downstream systems days later.
03What service levels should an agent have?
Availability and latency as usual, plus correctness or task success rate against an evaluation set, groundedness for retrieval-based answers, action-safety measured as policy violations per thousand actions, escalation rate and quality, and cost per task, each with a target and an error budget.
04What does agent observability require?
Traces at step granularity covering the prompt, retrieved context, model and version, tool calls with arguments and results, decisions, final output, and cost, joined to a request identifier so any interaction can be reconstructed completely after the fact.
05How does rollback work for an agent?
By versioning every component that changes behaviour, meaning prompts, models, tool definitions, retrieval indexes, and configuration, so any one can be reverted independently and quickly. Rollback that only covers application code leaves the most common causes of regression in place.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.