Checklist · 5 minute read
AI Agent Launch Checklist: What Must Be True Before Go-Live
Before an agent goes live it needs a written scope, least-privilege permissions with hard limits, an evaluation suite that passes, trajectory logging, a tested kill switch, a defined escalation path, and a named owner. Missing any one of these turns a manageable incident into an unmanageable one.
An agent that works in testing is not ready to act on real systems. This checklist covers what must be true before go-live, drawn from FISTA Solutions' AI agents production deployments.
What must be true before launch?
Seven gates, all of which should be closed.
| Gate | Evidence it is closed |
|---|---|
| Scope documented | A written scope approved by the owner |
| Permissions scoped | Tool list with limits, reviewed |
| Evaluation passing | Dated report on representative cases |
| Observability live | Trajectories visible in the log |
| Kill switch tested | A record of it being exercised |
| Escalation defined | A staffed path with response times |
| Owner named | A person, with authority |
Scope and boundaries
Write down what the agent does before deciding what it may access. An undocumented scope cannot be reviewed and expands silently.
- The agent's purpose is stated in one paragraph
- The cases it handles are listed explicitly
- The cases it must escalate are listed explicitly
- Actions requiring human approval are identified with thresholds
- The maximum consequence of a wrong action is written down
- Someone outside the build team has reviewed the scope
- The scope document is versioned alongside the code
Permissions and limits
Capability is the blast radius. Every tool the agent holds is something a failure can do, so the list should be as short as the task allows.
- Each tool is justified against a case in the scope
- Read and write capabilities are separated where possible
- Monetary limits per action and per period are enforced in code
- Volume limits per period are enforced in code
- The capability combination has been reviewed as a set, not tool by tool
- Authorisation is checked outside the model, not stated in the prompt
- Credentials are agent-specific and revocable independently
Evaluation
A passing evaluation suite is what distinguishes a tested system from a hopeful one. See how to build an agent evaluation harness.
- Representative cases drawn from real workload, not invented
- Known failure modes included as cases
- Edge conditions and malformed inputs included
- Adversarial inputs including injection attempts included
- Trajectories evaluated, not only final outputs
- Results dated, stored, and linked to the version tested
- The suite runs in the deployment pipeline and blocks on regression
Observability
If you cannot see what the agent did, you cannot operate it. See observability for web apps.
- Every step logged: input, reasoning, tool, arguments, result
- A correlation identifier links an entire task
- Model and prompt versions recorded with each action
- Cost per task attributed and visible
- Quality metrics monitored, not only availability
- Alerts defined on error rate, escalation rate, and cost anomalies
- Logs retained for a period that satisfies your retention policy
Failure handling and rollback
Production failures are certain. What matters is whether the behaviour when they occur was chosen or inherited.
- Behaviour defined for provider timeout, rate limit, and error
- Behaviour defined for tool failure mid-task
- Partially completed tasks leave a recoverable state
- The kill switch has been exercised in a rehearsal
- Actions taken by the agent can be identified and reversed where reversible
- Rollback to the previous agent version has been tested
- On-call knows how to disable it without an engineer
People and escalation
The exception path needs staffing before launch, not after the queue builds. See how to staff an AI support rotation.
- A named owner accountable for the agent's behaviour
- The owner has authority to disable it
- The escalation queue is staffed with a response expectation
- Escalation volume has been estimated and capacity provisioned
- Affected teams have been told what changes and when
- Customer-facing disclosure decided where the agent contacts people
- A rollback decision threshold is agreed in advance
What are the most common failures?
Launching without a tested kill switch. Permissions granted broadly because narrowing was inconvenient. Evaluation limited to happy paths. Logging final actions without reasoning. And an escalation path that exists on paper with nobody assigned.
Who should own this?
A named individual in the business function the agent serves, not in the engineering team that built it. The owner decides scope, accepts the residual risk, and holds authority to pause. Engineering owns the implementation; the business owns the behaviour.
How often should it run?
Before every launch, and again whenever scope expands, permissions change, or the underlying model changes. A quarterly re-run against deployed agents catches the drift that accumulates between deliberate changes.
What evidence should it produce?
A dated completed checklist, the scope document, the permission list with limits, the evaluation report, and a record of the kill switch rehearsal. That set answers most questions an auditor or an incident review will ask.
What about agents already in production?
Run the checklist retrospectively. Most deployed agents fail several items, and the common ones are trajectory logging, tested rollback, and a named owner.
Fix the kill switch and the owner first, because those bound the consequence of everything else. Then work through permissions, which is usually where the largest unnecessary exposure sits. See agent permission review checklist.
What should you do first?
Test your kill switch on the agent you already have running. If nobody has exercised it, that is the first item, and it takes an afternoon.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: agents launched against a written scope with permissions enforced in code, and kill switches rehearsed rather than documented, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI agent production readiness checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is the single most important item?
A tested kill switch. Everything else reduces the chance of a problem; the kill switch bounds how long a problem lasts, and it is the control on-call reaches for first.
02Why does scope need to be written?
Because an unwritten scope expands. A document stating what the agent may do, what it may not, and what requires approval is what makes permission review and incident assessment possible.
03What does trajectory logging add?
The ability to understand how a wrong action was reached. Logging only the final action tells you what happened and nothing about why, which makes investigation guesswork.
04How thorough should evaluation be before launch?
Representative cases, known failure modes, edge conditions, and adversarial inputs, all passing. A suite of happy-path examples predicts nothing about production behaviour.
05Can you launch without a named owner?
You can, and it is the most common gap. Without a named person with authority to disable the agent, an incident becomes a search for who decides while the agent continues acting.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.