FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Checklist ¡ 4 minute read

AI Escalation Matrix Template: Who Handles What, and When

An escalation matrix works when severity is defined by observable impact, routing names people rather than teams, response times are agreed with those people, and whoever receives an escalation has authority to act. Vague severity and unowned routes are why most matrices are ignored.

By FISTA Solutions¡ AI-Native Engineering Team¡
AI Escalation Matrix Template: Who Handles What, and When article cover

Escalation matrices fail when severity is vague and the route has no owner. This template covers what makes one function, drawn from FISTA Solutions' AI agents operational practice.

What does the matrix contain?

Six columns, filled per system.

ColumnWhat goes in it
SeverityObservable impact definition
TriggerWhat conditions produce this severity
First responderNamed rota with a paging route
Response timeAgreed, not aspirational
AuthorityWhat they may do without asking
Escalation from hereWho, and after how long

Severity definitions

Observable, not judged.

  • Each level defined by measurable impact
  • User numbers or percentages stated where relevant
  • Financial thresholds stated where relevant
  • Data exposure treated as its own category
  • Quality degradation given an explicit path
  • Examples given for each level
  • Ambiguous cases resolved upward by default

AI-specific triggers

Conventional definitions miss the failure that matters here.

  • Quality degradation thresholds defined and measurable
  • Agent taking incorrect actions given a severity
  • Cost anomaly thresholds defined
  • Provider outage and its customer impact mapped
  • Data exposure through model output given a severity
  • Retrieval failure or corpus problem given a path
  • Permission or limit breach given a severity

Routing

Named people with a reachable mechanism.

  • Each severity routes to a named rota
  • The paging or contact mechanism stated and tested
  • Backup contact named for each rota
  • Business owner contact for decisions beyond engineering
  • Communications contact for customer-facing incidents
  • Legal and security contacts for their categories
  • Vendor contacts with account details

Response times

Agreed by the people who will meet them.

  • Acknowledgement time defined per severity
  • Time to begin investigation defined
  • Update cadence during an incident defined
  • Times agreed by the responders, not imposed
  • Out-of-hours expectations stated explicitly
  • Coverage gaps stated honestly where they exist
  • Actual times measured against the commitment

Authority

The responder should be able to act.

  • Authority to disable a feature stated per role
  • Authority to roll back stated
  • Authority to disable an agent stated
  • Spend limits for remediation stated
  • Customer communication authority stated
  • Cases requiring a business decision identified in advance
  • Decision-maker reachable for those cases

Maintenance

A matrix naming someone who left is worse than none.

  • Owner named for the matrix itself
  • Reviewed quarterly
  • Reviewed after any incident with a routing failure
  • Contact details verified rather than assumed
  • Rota changes reflected promptly
  • Accessible during an outage, not only in a wiki that may be down
  • Tested by a drill at least annually

What are the most common failures?

Severity requiring judgement. Routing to team channels. Response times nobody agreed to. Responders without authority. And a matrix stored somewhere that is unavailable during an incident.

Who should own this?

Operations owns the matrix; each named responder's manager confirms the commitment; the system's business owner confirms the authority delegations.

How often should it run?

Quarterly review, plus after any incident where routing or response failed. Verify contact details each review rather than assuming they are current.

What evidence should it produce?

The current matrix with a review date, measured response times against the commitments, and drill records.

What if you have no out-of-hours coverage?

State that plainly rather than implying coverage that does not exist. A matrix promising a one-hour response overnight when nobody is on call is a commitment nobody can meet.

The honest alternative is a documented degraded state: the system disables itself or falls back automatically outside covered hours. See AI disaster recovery checklist.

What should you do first?

Check whether your current matrix names a person or a team. If it names teams, nobody is on call for it.

How FISTA Solutions helps

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: severity defined by observable impact, routing to named rotas with tested paging, and responders given authority to disable a system themselves, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To adapt this checklist to your environment, message FISTA on WhatsApp, or read how to staff an AI support rotation.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What makes severity definitions work?

Observable impact. Number of users affected, whether money moves, whether data is exposed. Definitions requiring judgement produce inconsistent severity and delayed response.

02Why route to people rather than teams?

Because a team is not on call. Routing to a named rota with a paging mechanism produces a response; routing to a team channel produces a message read the next morning.

03What is specific to AI incidents?

Quality degradation with no error. It needs its own severity path, because conventional definitions based on availability do not capture a system responding promptly with wrong answers.

04Why does authority matter?

Because an escalation reaching someone who must find a decision-maker adds the delay the escalation was meant to avoid. The receiver should be able to disable the system themselves.

05How often should it be reviewed?

After every incident where routing or response was wrong, and at least quarterly. People change roles and rotas change, and a matrix naming someone who left is worse than none.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project