FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Checklist · 5 minute read

AI Agent Launch Checklist: What Must Be True Before Go-Live

Before an agent goes live it needs a written scope, least-privilege permissions with hard limits, an evaluation suite that passes, trajectory logging, a tested kill switch, a defined escalation path, and a named owner. Missing any one of these turns a manageable incident into an unmanageable one.

By FISTA Solutions· AI-Native Engineering Team·
AI Agent Launch Checklist: What Must Be True Before Go-Live article cover

An agent that works in testing is not ready to act on real systems. This checklist covers what must be true before go-live, drawn from FISTA Solutions' AI agents production deployments.

What must be true before launch?

Seven gates, all of which should be closed.

GateEvidence it is closed
Scope documentedA written scope approved by the owner
Permissions scopedTool list with limits, reviewed
Evaluation passingDated report on representative cases
Observability liveTrajectories visible in the log
Kill switch testedA record of it being exercised
Escalation definedA staffed path with response times
Owner namedA person, with authority

Scope and boundaries

Write down what the agent does before deciding what it may access. An undocumented scope cannot be reviewed and expands silently.

  • The agent's purpose is stated in one paragraph
  • The cases it handles are listed explicitly
  • The cases it must escalate are listed explicitly
  • Actions requiring human approval are identified with thresholds
  • The maximum consequence of a wrong action is written down
  • Someone outside the build team has reviewed the scope
  • The scope document is versioned alongside the code

Permissions and limits

Capability is the blast radius. Every tool the agent holds is something a failure can do, so the list should be as short as the task allows.

  • Each tool is justified against a case in the scope
  • Read and write capabilities are separated where possible
  • Monetary limits per action and per period are enforced in code
  • Volume limits per period are enforced in code
  • The capability combination has been reviewed as a set, not tool by tool
  • Authorisation is checked outside the model, not stated in the prompt
  • Credentials are agent-specific and revocable independently

Evaluation

A passing evaluation suite is what distinguishes a tested system from a hopeful one. See how to build an agent evaluation harness.

  • Representative cases drawn from real workload, not invented
  • Known failure modes included as cases
  • Edge conditions and malformed inputs included
  • Adversarial inputs including injection attempts included
  • Trajectories evaluated, not only final outputs
  • Results dated, stored, and linked to the version tested
  • The suite runs in the deployment pipeline and blocks on regression

Observability

If you cannot see what the agent did, you cannot operate it. See observability for web apps.

  • Every step logged: input, reasoning, tool, arguments, result
  • A correlation identifier links an entire task
  • Model and prompt versions recorded with each action
  • Cost per task attributed and visible
  • Quality metrics monitored, not only availability
  • Alerts defined on error rate, escalation rate, and cost anomalies
  • Logs retained for a period that satisfies your retention policy

Failure handling and rollback

Production failures are certain. What matters is whether the behaviour when they occur was chosen or inherited.

  • Behaviour defined for provider timeout, rate limit, and error
  • Behaviour defined for tool failure mid-task
  • Partially completed tasks leave a recoverable state
  • The kill switch has been exercised in a rehearsal
  • Actions taken by the agent can be identified and reversed where reversible
  • Rollback to the previous agent version has been tested
  • On-call knows how to disable it without an engineer

People and escalation

The exception path needs staffing before launch, not after the queue builds. See how to staff an AI support rotation.

  • A named owner accountable for the agent's behaviour
  • The owner has authority to disable it
  • The escalation queue is staffed with a response expectation
  • Escalation volume has been estimated and capacity provisioned
  • Affected teams have been told what changes and when
  • Customer-facing disclosure decided where the agent contacts people
  • A rollback decision threshold is agreed in advance

What are the most common failures?

Launching without a tested kill switch. Permissions granted broadly because narrowing was inconvenient. Evaluation limited to happy paths. Logging final actions without reasoning. And an escalation path that exists on paper with nobody assigned.

Who should own this?

A named individual in the business function the agent serves, not in the engineering team that built it. The owner decides scope, accepts the residual risk, and holds authority to pause. Engineering owns the implementation; the business owns the behaviour.

How often should it run?

Before every launch, and again whenever scope expands, permissions change, or the underlying model changes. A quarterly re-run against deployed agents catches the drift that accumulates between deliberate changes.

What evidence should it produce?

A dated completed checklist, the scope document, the permission list with limits, the evaluation report, and a record of the kill switch rehearsal. That set answers most questions an auditor or an incident review will ask.

What about agents already in production?

Run the checklist retrospectively. Most deployed agents fail several items, and the common ones are trajectory logging, tested rollback, and a named owner.

Fix the kill switch and the owner first, because those bound the consequence of everything else. Then work through permissions, which is usually where the largest unnecessary exposure sits. See agent permission review checklist.

What should you do first?

Test your kill switch on the agent you already have running. If nobody has exercised it, that is the first item, and it takes an afternoon.

How FISTA Solutions helps

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: agents launched against a written scope with permissions enforced in code, and kill switches rehearsed rather than documented, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI agent production readiness checklist.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What is the single most important item?

A tested kill switch. Everything else reduces the chance of a problem; the kill switch bounds how long a problem lasts, and it is the control on-call reaches for first.

02Why does scope need to be written?

Because an unwritten scope expands. A document stating what the agent may do, what it may not, and what requires approval is what makes permission review and incident assessment possible.

03What does trajectory logging add?

The ability to understand how a wrong action was reached. Logging only the final action tells you what happened and nothing about why, which makes investigation guesswork.

04How thorough should evaluation be before launch?

Representative cases, known failure modes, edge conditions, and adversarial inputs, all passing. A suite of happy-path examples predicts nothing about production behaviour.

05Can you launch without a named owner?

You can, and it is the most common gap. Without a named person with authority to disable the agent, an incident becomes a search for who decides while the agent continues acting.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project