PagerDuty
Produit & IngénierieSix PagerDuty KPIs selected to measure on-call responsiveness and incident resolution quality per engineer, with the selection criteria made explicit.
Six PagerDuty KPIs selected to measure on-call responsiveness and incident resolution quality per engineer, with the selection criteria made explicit.
| Indicator | Object | Type | Formula | Unit |
|---|---|---|---|---|
| Mean Time to Acknowledge Average time between incident creation and first acknowledgment, per assignee. | Incident | Leading | AVG(time_to_acknowledge) | minutes |
| Mean Time to Resolve Average time between incident creation and resolution, per resolving engineer. | Incident | Lagging | AVG(time_to_resolve) | minutes |
| Incidents Resolved Number of incidents resolved, per resolving engineer. | Incident | Lagging | COUNT | count |
| P1/P2 Incidents Resolved Number of high-priority (P1 or P2) incidents resolved, per resolving engineer. | Incident | Lagging | COUNT | count |
| Escalation Rate Ratio of escalated incidents over total assigned incidents, per assignee. | Incident | Leading | COUNT_RATIO | % |
| Reopen Rate Ratio of incidents re-triggered within one hour of resolution over total resolved incidents, per resolving engineer. | Incident | Lagging | COUNT_RATIO | % |
PagerDuty exposes several object types: incidents, services, schedules, on-call shifts, escalation policies, teams, and log entries. This integration focuses exclusively on incidents, which represent the fundamental unit of on-call work and the only object type where individual performance is directly attributable. Service-level objects capture configuration and topology rather than individual contribution; schedule objects describe shift assignments without measuring what happens during those shifts. Six KPIs were retained, selected against three criteria: ability to attribute to an individual engineer, resistance to gaming, and balance between leading indicators of process discipline and lagging indicators of reliability outcomes.
Mean Time to Acknowledge measures the average elapsed time between an incident being created and the assigned engineer acknowledging it. It is the most direct measure of on-call responsiveness and the only KPI in this set that captures behavior in the first minutes of an incident. Mean Time to Resolve measures the average elapsed time between incident creation and resolution. These two indicators must be read in relation to each other. A short Mean Time to Acknowledge combined with a long Mean Time to Resolve indicates that engineers are engaging promptly but taking significant time to diagnose and fix, which may point to system complexity, tooling gaps, or knowledge deficits. The inverse pattern — a long acknowledgment time followed by a quick resolution — may suggest that engineers delay engaging until they are confident they can resolve immediately, a form of selective response that skews both metrics.
Incidents Resolved counts the total volume of incidents an engineer has closed. P1/P2 Incidents Resolved isolates the subset of critical-priority incidents within that total. The two indicators serve different analytical purposes: Incidents Resolved captures throughput, while P1/P2 Incidents Resolved captures contribution to the most consequential outages. An engineer with a high Incidents Resolved count but a low P1/P2 share may be handling a disproportionate volume of low-severity noise; an engineer with few total resolutions but a high critical share is carrying above-average systemic risk.
Reopen Rate addresses the primary gaming risk of Mean Time to Resolve. When an engineer resolves an incident prematurely — closing the ticket before the underlying issue is fixed — the system re-triggers the same incident within minutes. Reopen Rate measures this pattern directly: it counts incidents re-triggered within one hour of resolution as a proportion of total resolutions. A rising Reopen Rate paired with a falling Mean Time to Resolve is a reliable signal of superficial resolution behavior, not genuine improvement. Together, the two indicators make it analytically costly to optimize one at the expense of the other.
Escalation Rate measures the proportion of assigned incidents that an engineer escalates to the next level of the escalation policy rather than resolving independently. A low escalation rate generally indicates that the engineer is handling incidents within their scope of ownership, which is the intended design of an on-call rotation. A persistently high escalation rate for a given engineer may signal a skills gap, an under-specified runbook, or a mismatch between the services they are covering and their domain knowledge. Because escalations have visible costs — they wake up other engineers and extend resolution time — this indicator is structurally resistant to gaming: an engineer who avoids escalating when they should will generate other failures that surface in Mean Time to Resolve and Reopen Rate.
PagerDuty records incident lifecycle events and assignments, but does not measure the quality of the investigation or the permanence of the fix. A resolved incident may correspond to a thorough root-cause analysis or to a temporary workaround; the API data does not distinguish between them. The P1/P2 filter depends on the priority feature being enabled on the account and on teams applying priority classifications consistently: without disciplined tagging, P1/P2 Incidents Resolved will undercount or produce noise. Similarly, Mean Time to Acknowledge reflects the paging system working as configured — if engineers are not paged promptly, the metric measures infrastructure latency rather than individual responsiveness. The reliability of these six KPIs depends directly on teams maintaining clean incident data: consistent priority assignment, timely acknowledgment, and accurate resolution without premature closing.
No sign-up required and from your company's public data, we'll build a tailored scenario. Within a few hours, you'll receive an email with your access link.
We're preparing your personalized preview and will email you the link shortly.