FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Whitepaper · 9 minute read

Agent Reliability Engineering: Running AI Agents in Production

Agent reliability engineering is the practice of keeping AI agents dependable in production: defining service levels that include correctness, cataloguing failure modes that differ from ordinary software, running evaluation continuously rather than at release, instrumenting every step, bounding actions with guardrails and rollback, and operating a response model for failures that are silent rather than loud.

By FISTA Solutions· AI-Native Engineering Team·
Agent Reliability Engineering: Running AI Agents in Production article cover

Software reliability engineering matured around a specific assumption: systems fail visibly. They return errors, exceed timeouts, exhaust memory, or stop responding, and the discipline built its alerting, dashboards, and on-call practice around detecting and responding to those signals. AI agents break that assumption. An agent can return a well-formed response in two hundred milliseconds with a two hundred status code and be completely wrong, and no conventional monitor will notice. This whitepaper defines the practices that close that gap. It draws on FISTA Solutions' production operations across AI agents deployments and complements the AI observability whitepaper and ai agent observability.

What is agent reliability engineering?

The discipline of keeping agents dependable in production. It borrows the structure of site reliability engineering, meaning service levels, error budgets, observability, incident response, and blameless review, and extends it to the failure modes that only probabilistic, tool-using systems exhibit. Its central claim is that correctness is an operational property, not a release-time property, because agent behaviour changes when models, prompts, tools, retrieval content, or user inputs change, which is continuously.

What is the failure taxonomy?

Failure modeWhat happensWhy conventional monitoring misses it
Plausible incorrectnessFluent, confident, wrong answerReturns 200 with normal latency
Groundedness lossAnswer not supported by retrieved sourcesRetrieval succeeded; text looks cited
Retrieval degradationIndex stale or ranking degradedQuery returns results, just the wrong ones
Silent tool failureTool returns empty or partial data, agent proceedsTool call succeeded technically
Wrong actionAction within permissions but not appropriatePolicy allowed it
Loop and runawayAgent repeats steps, burning costOnly visible in cost, often late
Drift after provider changeModel updated, behaviour shiftedNothing in your code changed
Prompt injectionUntrusted content redirects the agentInput looked like ordinary data
Escalation failureAgent should have handed off and did notNo error; the user simply gave up
Cost blowoutContext growth or retries multiply spendLatency and errors normal

The first row is the defining one. In ordinary software, output that looks right usually is right. In agents, looking right is what the model optimises for, which inverts the relationship between appearance and correctness.

What service levels should an agent carry?

Availability and latency remain, but they are the least informative. The set that matters adds:

  • Task success rate: the proportion of interactions achieving the intended outcome, measured against an evaluation set and, where possible, verified downstream.
  • Groundedness: for retrieval-based answers, the proportion of claims supported by cited sources.
  • Action safety: policy violations or incorrect actions per thousand actions taken.
  • Escalation rate and quality: too low means the agent is trapping users; too high means it is not resolving.
  • Cost per task: with an alert threshold, because cost regressions are often the first visible sign of a behavioural regression.

Each gets a target and an error budget. When the budget is spent, changes stop and the team fixes quality before shipping features, which is the mechanism that keeps reliability from being traded away under delivery pressure. Service level design is in what is an slo for ai systems.

Why must evaluation run continuously?

Because the inputs change when nothing in your codebase does. Providers update models. Retrieval corpora are edited by people who do not know an agent depends on them. Upstream systems change their data formats. User behaviour shifts seasonally. A test suite that ran at release and passed says nothing about today.

Continuous evaluation means running a reference set on a schedule and on every change to any behaviour-affecting component, plus sampling live traffic for automated and human scoring. The reference set is a maintained asset with an owner, updated from production failures and retired when cases become obsolete. Practice is in what is continuous evaluation and what is a golden dataset.

What does observability require?

Traces at step granularity. For every interaction: the input, the resolved prompt, the retrieved context with document identifiers and scores, the model and exact version, each tool call with arguments and results, intermediate reasoning artifacts where available, the final output, the cost, and the latency of each step, all joined by a request identifier. Anything less makes post-hoc diagnosis guesswork, and agent diagnosis is almost entirely post-hoc because the failure was not visible when it happened.

On top of traces sit the aggregate signals: evaluation scores over time, groundedness rates, tool error and empty-result rates, escalation rates, cost distributions, and retrieval score distributions, which drift before quality does and therefore serve as a leading indicator. See what is distributed tracing for ai and how to build an agent trace analysis pipeline.

How do guardrails bound the damage?

By limiting what an agent can do, not by trying to make it always right. Permissions scoped to the agent's role. Action limits with value thresholds above which approval is required. Rate and volume caps that stop a loop before it becomes an incident. Output validation against schemas and business rules before anything is acted on. Confirmation steps for irreversible actions. And a kill switch that an on-call engineer can use without a deployment.

Blast radius is the design question: if this agent behaves as badly as it possibly could within its permissions, what is the worst outcome, and is that acceptable? Where the answer is no, the permissions are wrong. Guardrail design is in ai agent guardrails and permissions in how to design tool permissions for ai agents.

What does rollback cover?

Everything that changes behaviour, each versioned and revertible independently: prompts and system instructions, model selection and version pins, tool definitions and their schemas, retrieval indexes and their content snapshots, configuration and thresholds, and application code. A rollback capability that covers only code leaves the most common regression sources untouched, since most agent regressions come from prompt edits, provider model updates, or content changes.

Two practices make this real. Pin model versions rather than accepting provider defaults, so provider updates become a scheduled change with evaluation rather than a surprise. And snapshot retrieval indexes so a content change can be reverted as a unit. See what is a canary deployment and what is a shadow deployment.

How should changes be deployed?

Progressively, with evaluation gates. A behaviour-affecting change passes the reference evaluation, then runs in shadow against live traffic where the comparison is possible, then canaries to a small traffic share with quality metrics compared to the control, then expands. The gate is quality, not absence of errors.

For agents taking actions, shadow mode means executing the reasoning without the action and comparing proposed actions to those taken by the current system or by people. That catches action-selection regressions before they reach a customer account.

What does incident response look like?

Different from an outage. The detection is usually a metric drift, a sampled evaluation drop, a spike in escalations, or a customer report, not a page from a health check. The first response question is not "is it up" but "what changed", and the answer is usually a prompt edit, a provider model update, a content change, or an upstream data change.

Practical playbooks cover: quality regression, in which the team identifies the changed component, reverts it, and confirms via evaluation; wrong action taken, in which the team stops the agent, identifies affected records, remediates them, and only then investigates; prompt injection, in which the team isolates the content source and reviews what the agent accessed; and cost spike, in which the team caps and diagnoses whether it is volume or behaviour. Each ends with a blameless review that adds the case to the evaluation set, which is the mechanism that prevents recurrence. See ai incident response and the AI incident postmortem template.

What does on-call look like for agents?

The rotation needs people who can read a trace and judge whether an output was acceptable, which is a different skill from restarting a service. Runbooks specify how to disable an agent, how to revert each component, how to identify affected interactions, and who to notify when actions were taken in customer or financial systems.

Alerting should page on action-safety violations, cost anomalies, hard failures, and sustained evaluation score drops, and should ticket rather than page on softer quality signals, because the fatigue from paging on noisy quality metrics destroys the rotation faster than any other mistake. On-call design is in how to set up ai on call.

How does this scale across many agents?

Through shared infrastructure rather than per-agent heroics. A common gateway with version pinning and cost attribution, a shared tracing schema, a shared evaluation harness with per-agent reference sets, a shared guardrail library, and a single inventory recording every agent, its owner, its permissions, its service levels, and its current evaluation scores. Organisations running twenty agents without this end up with twenty different reliability practices and no ability to answer basic questions during an incident. See the agentic AI governance whitepaper.

What does maturity look like?

LevelDetectionEvaluationRollbackResponse
InitialCustomer reportsManual spot checksCode onlyAd hoc
ManagedBasic metrics, some alertsReference set at releaseCode and promptsRunbook exists
DefinedTraces plus quality metricsContinuous plus samplingAll components versionedPlaybooks per failure type
OptimisedLeading indicators, drift alertsGated deployment, shadow, canaryAutomated, tested regularlyBlameless review feeding evaluation

Most organisations sit between initial and managed while believing they are at defined, because their agents have not yet failed in a way they noticed.

How FISTA Solutions delivers this

FISTA Solutions builds and operates production agents with service levels that include correctness, step-level tracing, continuous evaluation, versioned rollback across prompts, models, tools, and indexes, and incident playbooks written for silent failures, delivered through AI enablement, AI agents, and forward deployed engineers who set up the reliability practice inside client teams. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To run agents you can depend on, message FISTA on WhatsApp, or read the AI observability whitepaper.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What is agent reliability engineering?

The operational discipline of keeping AI agents dependable in production, combining service level definition that includes correctness, failure mode analysis, continuous evaluation, step-level observability, guardrails and rollback, and incident response designed for silent quality failures rather than outages alone.

02How do agent failures differ from software failures?

Ordinary software fails loudly: errors, timeouts, crashes. Agents often succeed technically while being wrong, producing fluent, plausible output or taking a defensible-looking wrong action. Nothing alerts, dashboards stay green, and the failure surfaces through customers or downstream systems days later.

03What service levels should an agent have?

Availability and latency as usual, plus correctness or task success rate against an evaluation set, groundedness for retrieval-based answers, action-safety measured as policy violations per thousand actions, escalation rate and quality, and cost per task, each with a target and an error budget.

04What does agent observability require?

Traces at step granularity covering the prompt, retrieved context, model and version, tool calls with arguments and results, decisions, final output, and cost, joined to a request identifier so any interaction can be reconstructed completely after the fact.

05How does rollback work for an agent?

By versioning every component that changes behaviour, meaning prompts, models, tool definitions, retrieval indexes, and configuration, so any one can be reverted independently and quickly. Rollback that only covers application code leaves the most common causes of regression in place.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project