Playbook ¡ 6 minute read
How to Run a Production Readiness Review for AI Systems
A production readiness review for an AI system checks evaluation, monitoring, escalation, authority scope, cost limits, and the rollback path â none of which conventional service checklists cover. Findings in those areas should block launch rather than become follow-up items.
Conventional production readiness checklists cover availability, capacity, and security. None of them catches a system that stays up while producing confidently wrong results. This playbook covers what an AI readiness review should examine, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
Before any AI system serves real users or takes real actions, and again before any significant expansion of its audience or authority.
It is also worth running retrospectively on systems already in production that never had one. Those reviews reliably find at least one blocking gap, which is uncomfortable and better known than not.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Evaluation | Exists, ran recently, covers hard cases |
| 2. Quality monitoring | Detects drift, not only outages |
| 3. Escalation | Reaches someone who can resolve |
| 4. Authority scope | Irreversible actions gated |
| 5. Limits | Cost, steps, rate, circuit breakers |
| 6. Rollback | Tested, not documented |
Step 1 â Check evaluation exists and is current
Ask to see the evaluation set, the rubric, the most recent run, and the date.
A suite that exists and last ran two months ago describes a system that has since changed. A suite covering only cases the system handles well is a formality. A suite with no agreed rubric produces numbers nobody can interpret.
The question that settles it: what does the team do when the score drops. If the answer is not specific, evaluation is a report rather than a control. See how to run an ai evaluation program.
Step 2 â Check monitoring covers quality
Availability monitoring is necessary and insufficient. An AI system degrades by producing worse output, not by returning errors.
Look for output quality sampling, drift detection on inputs, escalation and override rates, and alerts with an owner. Ask what would fire if the model started producing subtly worse answers tomorrow â if the answer is a customer complaint, that is a blocking finding.
Check the alerts are routed to someone who will act rather than to a channel nobody watches.
Step 3 â Check escalation reaches someone useful
Test it, in the review, with a real case.
The common failure is an escalation path that routes to a queue with no context, or to someone who can only repeat what the system said. Both produce a worse experience than no automation, and both are invisible until a user hits them.
Ask what the person receiving an escalation sees, what authority they have, and how quickly they respond. See what is an escalation policy.
Step 4 â Check authority scope
Ask for the list of actions the system can take and which are gated.
Irreversible, customer-visible, and costly actions should require approval or be outside the system's reach entirely. An agent that can issue a refund, send an email, or change a record without a gate is a risk decision someone should have made explicitly.
Check the tools enforce it rather than the prompt. Instructions not to do something are not a control. See what is least privilege for ai agents.
Step 5 â Check the limits
Maximum steps per task, cost ceiling per task and per period, rate limits, and circuit breakers on repeated identical calls.
Agents loop, and loops are billed. A system without a step limit can consume a month's budget in an afternoon, and the first anyone knows is the invoice.
Ask what happens when a limit is hit: a clean failure with an escalation is acceptable, a silent partial result is not.
Step 6 â Check the rollback has been tested
Not documented â tested. Ask when it was last exercised and by whom.
A rollback plan that has never been run is a hypothesis. Configurations drift, dependencies change, and the moment you need it is the worst time to discover the feature flag no longer disables what it used to.
Exercise it during the review if it has not been tested recently. Fifteen minutes then is cheaper than an hour during an incident.
What about the data and privacy checks?
Cover them explicitly rather than assuming the general security review did.
What data reaches the model, where it is processed, what the provider retains, whether prompts and outputs are logged and for how long, and whether the system respects access controls when retrieving. Those are AI-specific questions a conventional review will not ask.
Access control at retrieval time is the one most often missed, and it fails as a security incident rather than a quality problem.
Who owns the system after launch?
Ask, and get a name.
Systems without a named production owner degrade quietly because nobody's job includes noticing. That owner needs to be someone who will see the monitoring, act on the alerts, and be accountable for the quality number.
If nobody will take it, that is a blocking finding. An unowned AI system in production is a liability with a launch date.
Who needs to be involved?
The engineering owner, an operator from the function that will run it, a reviewer who did not build it, and whoever will carry the pager.
The outside reviewer is what makes the review useful. Teams reviewing their own work confirm what they already believe.
How long does it take?
One to two hours for a prepared system, plus whatever remediation the findings require. Reviews taking a full day are finding problems that should have surfaced during the build.
What are the common failure modes?
Using a conventional checklist. Accepting evaluation that has not run recently. Not testing escalation. Gating by prompt rather than by tool. No limits. And a rollback nobody has exercised.
How do you know it worked?
No blocking findings at launch, the on-call person confident about what to do when it misbehaves, and the first month passing without a surprise that the review should have caught.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Ask the team what would alert them if output quality dropped by a third tomorrow. The answer tells you most of what the review will find.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: reviews that test escalation and rollback rather than reading about them, blocking findings treated as blocking rather than converted to tickets, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to roll out an AI agent to production.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why do AI systems need a different review?
Because the failure modes differ. Conventional checklists cover availability, capacity, and security, and none of them catch a system that stays up while producing confidently wrong results at scale.
02What should block a launch?
No evaluation, no quality monitoring, no working escalation path, unscoped authority on irreversible actions, absent cost or step limits, and an untested rollback. Each of those makes a bad outcome unrecoverable rather than merely possible.
03Who should be in the review?
The engineering owner, someone from the function that will operate it, a reviewer who did not build it, and whoever will be on call. Reviews run only by builders miss what operators will face.
04How long should it take?
An hour or two for a prepared system. Reviews that take a day are discovering problems that should have been found earlier, which is useful information about the process as well as the system.
05What if a finding cannot be resolved before launch?
Then either the launch moves or the scope shrinks. Launching with a known gap and a follow-up ticket is how systems reach production without escalation paths, and those tickets rarely get done.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.