Playbook ¡ 6 minute read
How to Staff an AI Support Rotation That Is Sustainable
An AI support rotation differs from conventional on-call because the failures are quality problems rather than outages. It needs runbooks covering quality incidents, people who can judge output as well as read logs, alerts that mean something, and a load that does not burn people out.
AI systems need support that differs from conventional on-call, because the characteristic failure is a system that is up and producing worse results. This playbook covers staffing for that, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
Before an AI system serving real users reaches production. Systems without a defined support arrangement get supported by whoever happens to notice, which is unsustainable and unfair.
It is also worth revisiting when a system's audience grows materially, because the volume and severity of incidents change together.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Define what it covers | Quality as well as availability |
| 2. Choose who can serve | Technical and domain judgement |
| 3. Write the runbooks | For the failures that actually happen |
| 4. Fix alert quality | Every page should need action |
| 5. Grant response authority | Act first, approve after |
| 6. Measure and rebalance | Before people leave |
Step 1 â Define what the rotation covers
Name the incident types: quality degradation, provider outage or rate limiting, unsafe or inappropriate output, cost spike, escalation overload, and conventional availability failures.
Quality degradation is the one conventional arrangements miss. A system returning responses within latency targets while producing materially worse answers is an incident, and nothing in a standard on-call setup detects or handles it.
Write the list down and agree severity levels. Without that, everything is either urgent or ignored.
Step 2 â Choose who can serve
Responders need to diagnose technically and judge whether output is acceptable in the domain.
Engineers alone miss domain-specific wrongness â an answer that reads well and is clinically or commercially wrong. Domain people alone cannot trace a trajectory or check a provider status page.
The practical answers are cross-training, pairing an engineer with a domain contact who can be reached, or a two-tier rotation. Whichever you choose, name it rather than assuming the engineer will know.
Step 3 â Write runbooks for real failures
For each incident type: how it is detected, what to check first, immediate mitigation, who to involve, and what to record.
Quality degradation needs a specific runbook: check whether the model version changed, whether the retrieval corpus changed, whether input distribution shifted, and whether a recent deployment correlates.
Runbooks written from imagined failures age badly. Update them after each incident with what was actually useful, and delete steps that nobody used. See how to build an ai runbook.
Step 4 â Fix alert quality relentlessly
Every page should require action. Alerts that fire without needing a response train people to ignore the pager, which is how real incidents get missed.
Review alerts regularly: how many fired, how many needed action, and which ones fire repeatedly for the same reason. Non-actionable alerts should be made actionable or deleted, not tolerated.
This is the single largest determinant of whether a rotation is sustainable, more than volume or hours.
Step 5 â Grant authority to act
Responders should be able to disable a feature, route to a fallback, restrict the system's authority, or roll back a change without seeking approval.
An on-call engineer who must wake a manager to disable a misbehaving agent cannot respond quickly, and the delay is the difference between a contained incident and a serious one.
Write the authority down, including its limits. Ambiguity produces hesitation at exactly the wrong moment.
Step 6 â Measure load and rebalance
Track pages per shift, out-of-hours pages, time spent, and how often the same issue recurs.
Recurring issues are a backlog problem presenting as a rotation problem. Fixing the top recurring cause usually reduces load more than adding people does.
Rebalance before people burn out rather than after. The signal that a rotation is unsustainable usually arrives as a resignation, at which point the load falls on fewer people.
Can the rotation be shared with existing on-call?
Often, and with adjustments. Adding AI systems to an existing rotation works where responders have the judgement and the runbooks exist.
What does not work is adding them silently. An on-call engineer paged about an AI quality issue with no runbook and no domain contact is being set up to fail, and the resulting experience poisons the arrangement.
What about offshore or follow-the-sun coverage?
It works well and depends entirely on handover quality. A shift handing over with a one-line note effectively restarts every open incident.
Require a structured handover: open incidents, what was tried, what is pending, and what to watch. Time zone coverage is one of the genuine advantages of a distributed team, and handover discipline is what realises it.
Who needs to be involved?
Engineers who can diagnose, domain contacts who can judge, and a manager who owns the rotation's sustainability.
The last role is frequently missing. Rotations without an owner accumulate load until someone leaves.
How long does it take?
Two to three weeks to define, write runbooks, and fix the worst alerts. Sustainability is continuous rather than a setup task.
What are the common failure modes?
Treating AI incidents as conventional outages. Responders without domain judgement. Runbooks for imagined failures. Tolerating non-actionable alerts. Responders without authority. And rebalancing after a resignation.
How do you know it worked?
Incidents detected by monitoring rather than by users, responders acting without escalation for the common cases, non-actionable alert volume falling, and rotation load stable across quarters.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Review the last month of alerts and count how many required action. That percentage tells you whether your rotation is currently sustainable.
What should be handed to the rotation at launch?
A runbook, a dashboard, an escalation contact, and a named owner who is not on the rotation. Systems handed over without those get supported by whoever built them, permanently, which is how teams end up unable to start anything new.
Make the handover a gate rather than a formality: the rotation should be able to decline a system that arrives without them.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: runbooks covering quality degradation rather than only availability, alert quality reviewed so every page needs a response, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to set up AI on-call.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How is AI support different from conventional on-call?
The failures differ. Conventional on-call handles outages and errors; AI support handles a system that is up and producing worse results, which requires judgement about output quality rather than log reading alone.
02Who should be on the rotation?
People who can both diagnose technically and judge whether output is acceptable. Pure engineers miss domain-specific wrongness; pure domain people cannot diagnose. Pairing or cross-training addresses it.
03What should runbooks cover?
Quality degradation, provider outages and rate limits, unsafe outputs, cost spikes, and escalation overload â each with detection, immediate action, and who to involve. Conventional runbooks cover none of these.
04What authority does a responder need?
To disable a feature, route to a fallback, restrict the system's authority, or roll back a change, without seeking approval. Responders who must wake someone to act cannot act quickly.
05How do you keep it sustainable?
By fixing alert quality relentlessly. A rotation paged for things that need no action burns people out regardless of volume, and every non-actionable alert should either be made actionable or removed.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.