Playbook ¡ 6 minute read
How to Set Up AI On-Call That People Can Sustain
On-call for AI systems needs a scope that includes quality failures, responders who can judge output as well as read logs, authority to disable and roll back without approval, and alert discipline strict enough that every page means something actionable.
On-call for AI systems differs from conventional on-call in scope, in who can respond, and in what authority they need. This playbook covers setting up a rotation people can sustain, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
Before an AI system serving real users reaches production. Systems without a rotation are covered by whoever happens to notice, which is unreliable and unfair to the person who ends up noticing.
It is also worth revisiting when a system's audience grows, because volume and severity change together.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Define what pages | Quality as well as availability |
| 2. Choose who can serve | Technical plus domain judgement |
| 3. Grant authority | Act first, approve after |
| 4. Attach runbooks | Every alert, a linked entry |
| 5. Enforce alert discipline | Actionable or removed |
| 6. Manage handover and load | Both determine sustainability |
Step 1 â Define what pages someone
Unsafe or inappropriate output, customer-visible failure, quality below threshold, cost spiking, and availability failures.
Quality below threshold is the one conventional setups omit. A system that is up and producing materially worse answers is an incident, and if nothing pages for it the first notification will be a complaint.
Everything else â a single odd output, a minor cost variance, a slow trend â belongs in the daily check rather than on the pager.
Step 2 â Choose who can serve
Responders need to diagnose technically and judge whether output is acceptable in the domain.
An engineer alone will miss domain-specific wrongness: an answer that reads well and is clinically, legally, or commercially incorrect. A domain expert alone cannot trace a trajectory or check a provider status page.
The practical arrangements are cross-training, a reachable domain contact, or a two-tier rotation. Choose one explicitly rather than assuming the engineer will manage.
Step 3 â Grant authority to act
Disable a feature, roll back a prompt or model change, route to a fallback, restrict the agent's authority â without seeking approval.
An on-call engineer who must wake a manager to disable a misbehaving system cannot respond quickly, and the delay is the difference between containment and a larger incident.
Write the authority down including its limits, so there is no hesitation at the moment it matters. Ambiguity produces a phone call instead of an action.
Step 4 â Attach runbooks to every alert
Each alert carries a link to its entry: what it means, what to check, what to do.
An alert without a runbook is a notification that something is wrong with no guidance, which produces improvisation by whoever is least equipped.
If writing the entry proves difficult, the alert is probably not well defined. That is a useful signal about the alert rather than about the runbook. See how to build an ai runbook.
Step 5 â Enforce alert discipline
Review alerts regularly: what fired, what needed action, what was noise. Fix or remove everything in the third category.
This is the single largest determinant of whether a rotation is sustainable, ahead of volume and hours. A pager that fires for things needing no response trains people to ignore it, and the alert that mattered arrives in the same stream.
Track the proportion of pages that led to action. Below about half, the alerting is the problem rather than the system.
Step 6 â Manage handover and load
Structured handover â open incidents, what was tried, what is pending, what to watch â and a regular look at pages per shift and out-of-hours frequency.
Follow-the-sun coverage is one of the genuine advantages of a distributed team and it lives or dies on handover. A one-line note restarts every open incident.
Rebalance before someone resigns. A rotation's unsustainability is usually first reported as a departure, at which point the load falls on fewer people. See how to staff an ai support rotation.
What is the handover format?
Open incidents with their current state, anything degraded but not incident-level, changes made during the shift, and anything to watch.
Written, structured, and short. Handovers conducted by conversation lose detail and cannot be referred back to; handovers written as prose take too long to read at the start of a shift.
A template takes five minutes to fill in and saves the incoming shift twenty.
Should AI systems share the existing rotation?
Frequently yes, with adjustments: runbooks for the AI-specific incidents, a domain contact available, and the responders briefed on what they are now covering.
What fails is adding them silently. An on-call engineer paged about a quality issue with no runbook and no domain contact has been set up to fail, and the experience sours the arrangement for everyone.
Who needs to be involved?
The engineers on the rotation, domain contacts who can judge output, and a manager accountable for the rotation's sustainability.
The last role is the one most often absent. Rotations without an owner accumulate load until it becomes visible as attrition.
How long does it take?
Two to three weeks to define scope, write runbooks, and tune the worst alerts. Sustainability is continuous rather than a setup task.
What are the common failure modes?
Paging only for outages. Engineers without domain support. No authority to act. Alerts without runbooks. Tolerating noise. And handover by conversation.
How do you know it worked?
Incidents detected by alerts rather than users, responders acting without escalation for common cases, alert noise falling, and rotation load stable across quarters.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Check what would page someone if output quality dropped by a third tomorrow. In most organisations the answer is nothing, and that is the gap to close first.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: paging scope that includes quality failures, responders granted the authority to disable and roll back without seeking approval, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to build an AI runbook.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What should page someone?
Unsafe output, customer-visible failure, quality dropping below threshold, cost spiking, and conventional availability failures. Anything that can wait until the morning check should not page, because tolerated noise destroys the pager's meaning.
02Who can serve on the rotation?
People who can diagnose technically and judge whether output is acceptable, or an engineer paired with a reachable domain contact. Pure engineering rotations miss domain-specific wrongness entirely, which is the failure that matters.
03What authority do responders need?
To disable a feature, roll back a change, route to a fallback, or restrict the system's authority, without seeking approval. Responders who must wake a manager to act cannot act quickly enough to contain anything.
04How do you keep it sustainable?
By removing or fixing every alert that fires without needing action. Alert quality is the main determinant of whether a rotation burns people out, ahead of volume or out-of-hours frequency.
05What about follow-the-sun coverage?
It works and depends entirely on handover quality. A shift handing over with a one-line note effectively restarts every open incident, which the customer experiences as being asked the same questions twice.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.