Playbook · 5 minute read
How to Build a PagerDuty AI Agent
A PagerDuty AI agent enriches incidents as they open with a responder briefing drawn from monitoring, recent changes, and similar past incidents, drafts stakeholder updates for a person to send, and assembles postmortem timelines. It does not page, escalate, reassign, or resolve incidents, because those are commitments between people that the agent cannot see.
A page at three in the morning arrives with a service name, an alert title, and nothing else. The responder then spends the first ten minutes discovering which service this is, what depends on it, whether anything was deployed, whether this happened before, and where the runbook is. An agent that does that work in the seconds between the alert firing and the phone buzzing is worth more than most observability investments, because it removes the worst part of on-call. This guide covers building one that helps without ever paging anyone itself, drawing on FISTA Solutions' AI agents delivery in operations. It complements the AI incident management whitepaper and how to build a datadog ai agent.
Where is the boundary?
| Activity | Verdict | Reason |
|---|---|---|
| Enrich an incident with a briefing note | Yes | Adds context, changes nothing |
| Draft a stakeholder update | Yes, for review | Tone under pressure is human |
| Suggest a likely cause with evidence | Yes, labelled | Responder confirms |
| Assemble postmortem timeline | Yes | Tedious and factual |
| Trigger a page | No | Human cost, trust risk |
| Escalate or reassign | No | Judgement about people |
| Acknowledge or resolve | No | Commitments between people |
| Change escalation policies or schedules | No | Organisational decisions |
The agent informs the people PagerDuty's escalation policies already chose. It does not choose them.
How does integration work?
Through webhooks on incident lifecycle events, which fire when an incident triggers, is acknowledged, escalates, or resolves, and the REST API for reading incident details, the affected service and its dependencies, past incidents on that service, notes, and runbook links. The API key is scoped to read plus the ability to add notes, and it holds nothing that can trigger, escalate, reassign, or resolve.
What goes in a responder briefing?
The alert and what it means in plain language. The affected service, its owners, and its upstream and downstream dependencies. Deployments and configuration changes on that service and its dependencies in the preceding hours. Other incidents currently open on related services. Similar past incidents on this service, with what the resolution turned out to be. The runbook link if one exists. And where the monitoring integration allows, the metric shape and top error patterns around the alert window.
All of that is attached as a single note within seconds of the incident opening, so it is on the responder's screen when they acknowledge. See how to build a datadog ai agent for the monitoring side of the briefing.
Why are similar past incidents the most useful context?
Because most incidents have happened before. The service that pages on the first of every month, the dependency that times out under a particular load pattern, the deploy that always needs a cache flush afterwards: experienced responders know these, and new ones do not. An agent that finds the three most similar past incidents on this service and reports how each was resolved transfers that knowledge at the moment it is needed.
Similarity should weight the same service, the same alert type, and the same error patterns, and the past incident's resolution notes are the payload. This is also the strongest argument for writing good resolution notes.
How should stakeholder updates be handled?
Drafted, not sent. During an incident someone must tell customers, executives, or other teams what is happening, and drafting that under pressure is slow and error-prone. The agent drafts an update from the current incident state, in the organisation's standard format and tone, and the incident commander edits and sends it.
The agent does not post to status pages or stakeholder channels directly. A misjudged update during an outage is a second incident, and tone under pressure is a human judgement.
How does postmortem assembly work?
The agent reads the incident's full record of state changes, notes, responder actions, and timings, correlates it with deployment and monitoring events, and drafts the chronology section of the postmortem with timestamps and sources. That is the part of every review that someone reconstructs by hand from scattered records, usually days later, usually incompletely.
The analysis sections, meaning contributing factors, what went well, and actions, remain the team's work, because they involve judgement and blameless discussion the agent should not pre-empt. See the AI incident postmortem template.
How is noise avoided?
One briefing note per incident, attached at open. Updates to the note only when something material changes, such as a related incident opening. Drafts for people rather than posts to channels. No new incidents created by the agent. An agent that comments continuously is muted by the second week, and a muted agent helps nobody.
How is it evaluated?
By responders, after incidents. Did the briefing contain what they needed. Did the similar incidents include the one that mattered. Was the suggested cause right, wrong, or misleading. Track time from acknowledgement to first meaningful action before and after, which is the outcome the briefing is meant to move. And confirm by key audit that the agent has never paged, escalated, or resolved.
What does the build sequence look like?
One week on webhooks, scoped keys, and the note-writing path. Two weeks on the briefing with service dependency and change data, reviewed by on-call engineers. One week on similar incident retrieval. One week on stakeholder update drafting. Postmortem assembly last, since it depends on the note quality the earlier stages establish.
What goes wrong?
Keys that can page. Agents that resolve incidents because the alert cleared. Updates posted directly to status pages. Continuous commentary that gets muted. Briefings that are complete and unreadable at three in the morning, which is a formatting problem. And similar-incident retrieval that surfaces incidents with the same title and nothing else in common.
How FISTA Solutions helps
FISTA Solutions builds PagerDuty agents that brief responders at incident open with dependencies, changes, and similar past incidents, draft stakeholder updates for the commander to send, assemble postmortem timelines, and hold no ability to page, escalate, or resolve, through AI enablement, AI agents, and forward deployed engineers working with SRE teams. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To make the first ten minutes of every incident shorter, message FISTA on WhatsApp, or read the AI incident management whitepaper.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How does an agent integrate with PagerDuty?
Through webhooks that fire on incident lifecycle events such as trigger, acknowledge, and resolve, and the REST API for reading incident details, services, past incidents, and notes, with an API key scoped to read plus note-writing and nothing that can page, escalate, or resolve.
02What is a responder briefing?
A note attached to the incident within seconds of it opening that summarises the alert, the affected service and its dependencies, recent deployments and changes, related open incidents, similar past incidents with what resolved them, and the runbook if one exists, so the responder starts informed.
03Why should the agent not page or escalate?
Because paging someone is a decision with human cost, escalation reflects judgement about severity and availability the agent lacks, and an agent that pages wrongly at three in the morning destroys trust in the whole system. Escalation policies already encode the organisation's decisions; the agent informs them.
04How does postmortem assembly work?
By reading the incident's timeline of state changes, notes, and responder actions alongside monitoring and deployment events, and drafting the chronology section of the postmortem with timestamps and sources, which is the part reviewers most often reconstruct by hand from scattered records.
05How is the agent kept from adding noise?
By writing one briefing note per incident rather than continuous commentary, by enriching incidents that already exist rather than creating new ones, and by drafting updates for a person to review rather than posting to stakeholder channels directly.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.