On-Call and Incident Response Playbook
A practical playbook for on-call engineers: triage, escalation, communication, and postmortems. Reduce MTTR and build a resilient incident response culture.
Note: This guide follows English-language naming conventions and terminology standards common in international development teams. Examples use English identifiers and comments to maximize compatibility across codebases and tooling.
On-Call and Incident Response Playbook
Introduction
Incidents are inevitable. What separates resilient teams from fragile ones is not the absence of failures, but the speed and quality of their response. This playbook provides a structured approach to handling production incidents — from the first alert to the postmortem.
The Incident Response Lifecycle
Detect → Triage → Mitigate → Resolve → Postmortem
↑ │
└────────── Monitor & Communicate ─────────┘
1. Detection
Alerting Principles
| Alert | Why It Matters | Threshold |
|---|---|---|
| Error rate spike | Users are seeing failures | > 0.1% of requests for 2 minutes |
| Latency p99 | Degraded user experience | > 500ms for 5 minutes |
| Saturation | Resource exhaustion approaching | CPU > 80%, memory > 85%, disk > 90% |
| Dependency failure | Downstream service is down | Health check fails 3 times |
Alert Fatigue Is Real
If an alert fires and the on-call engineer does not take action, it is not an alert — it is noise. Remove or downgrade alerts with > 80% false positive rate.
2. Triage
The FIRST Minute Checklist
When paged, answer these questions in order:
- What is failing? — service name, endpoint, region
- Who is affected? — all users, a subset, internal only?
- When did it start? — exact time of first failure (check deployment logs)
- What changed? — any deploy, config change, or dependency shift?
- Is it getting worse? — trend of error rate over time
Severity Levels
| Severity | Definition | Response Time | Example |
|---|---|---|---|
| SEV-1 | Complete service outage or data loss | 15 minutes | Payment system down for all users |
| SEV-2 | Major functionality degraded | 30 minutes | Search returns empty for 50% of users |
| SEV-3 | Minor impact or workaround exists | 2 hours | Admin dashboard slow, API still fast |
| SEV-4 | No user impact, potential risk | Next business day | Log volume spike, no errors yet |
3. Mitigation
Stop the Bleeding First
Your first goal is not to fix the root cause — it is to restore service. Prefer rollback over forward-fix during an incident.
# Rollback a bad deployment
kubectl rollout undo deployment/api-service
# Enable a feature flag kill switch
curl -X POST "https://config-service/flags/checkout-v2" \
-d '{"enabled": false}'
# Scale up to absorb load
kubectl scale deployment/api-service --replicas=20
Common Mitigation Tactics
| Problem | Fast Mitigation |
|---|---|
| Bad deployment | Rollback to last known good version |
| Traffic spike | Scale horizontally, enable rate limiting |
| Dependency failure | Enable circuit breaker, serve stale cache |
| Database overload | Kill slow queries, add read replicas |
| Configuration error | Revert config, restart with previous values |
4. Communication
Internal Status Updates
Post in your incident channel every 10 minutes:
[SEV-2] Checkout latency elevated
- Started: 14:32 UTC
- Impact: ~30% of checkout requests timeout
- Cause: database connection pool exhausted after v2.4.1 deploy
- Mitigation: rolled back to v2.4.0 at 14:45, monitoring recovery
- ETA: 15:00 UTC if trend holds
- Commander: @alice
External Communication
| Severity | External Notice? | Who |
|---|---|---|
| SEV-1 | Yes, immediate | Customer support + status page |
| SEV-2 | Yes, if > 30 min | Customer support + status page |
| SEV-3 | No, unless asked | Internal only |
| SEV-4 | No | Internal only |
Blameless Communication Rules
- Do not name individuals as causes
- Do not use “human error” as a root cause
- Focus on what happened, what was done, and what is next
5. Resolution
Definition of Resolved
An incident is resolved when:
- Error rates return to baseline for 10 minutes
- All mitigations are stable
- No new symptoms have appeared
- The incident commander declares “all clear”
After All Clear
- Stop the clock (log total incident duration)
- Schedule postmortem within 24 hours for SEV-1/2
- Create follow-up tickets with owners and due dates. Update CI/CD if needed.
- Update runbooks with anything learned
6. Postmortem
The Five Whys
Ask “why” recursively until you reach a systemic issue, not a symptom.
Problem: Payment API returned 500 errors for 20 minutes.
Why? → Database connection pool was exhausted.
Why? → v2.4.1 increased default pool size but forgot to close connections in new retry logic.
Why? → The change was not tested under load.
Why? → Load tests do not cover the checkout flow.
Why? → Load test scenarios were last updated 6 months ago.
Action: Add checkout flow to weekly [load tests](/recipes/performance/load-testing-k6); require load test pass in [CI](/guides/devops/cicd-pipeline-guide).
Postmortem Template
# Postmortem: [Incident Name] ([SEV-X])
## Summary
- Date: 2024-06-12
- Duration: 23 minutes
- Impact: 12% of checkout attempts failed
## Timeline
- 14:32 — First alert: error rate spike on /api/checkout
- 14:35 — On-call acknowledged
- 14:40 — Identified connection pool exhaustion
- 14:45 — Rolled back to v2.4.0
- 14:55 — Error rates returned to baseline
## Root Cause
v2.4.1 introduced a retry loop that leaked database connections.
## What Went Well
- Rollback completed in under 5 minutes
- Monitoring clearly pointed to connection pool exhaustion
## What Went Wrong
- Load tests did not cover the new retry logic
- No connection leak detection in staging
## Action Items
| Action | Owner | Due Date |
|--------|-------|----------|
| Add checkout flow to load tests | @bob | 2024-06-19 |
| Add connection leak alert | @alice | 2024-06-15 |
What Works
- Rotate on-call fairly — no one should be on-call more than 1 week in 4
- Compensate for off-hours — pay extra or give time off in lieu
- Shadow on-call — new engineers shadow for 2-4 weeks before taking the pager
- Automate runbooks — if a runbook step is manual, add it to your automation backlog
- Review alerts quarterly — remove noise, tune thresholds, fix flapping alerts
Common Mistakes
- Skipping postmortems because “we are too busy”
- Blaming individuals instead of fixing systems
- Forward-fixing during an incident instead of rolling back
- Communicating too late to customers
- Not having a secondary on-call for escalation
- Keeping the same person on-call for weeks
Frequently Asked Questions
What if I do not know how to fix the issue?
That is expected. Your job is to contain the impact and find the right person — not to know every system. Escalate early and clearly. A 5-minute escalation is better than a 30-minute solo struggle.
How do I balance incident response with feature work?
Incidents are unplanned work. Track them. If a team spends > 20% of sprint capacity on incidents, that is a signal to invest in reliability (tests, automation, refactoring) rather than new capabilities.
Should junior engineers be on-call?
Yes, with mentorship. Shadowing senior engineers during incidents is one of the fastest ways to learn how systems fail. Start with low-severity rotations and pair them with a senior for the first month.
Advanced Topics
Scenario: Sev-1 Incident Response in E-commerce
Incident: Checkout down, 100% users affected
Severity: Sev-1 (critical)
Start: 14:12 UTC
On-call: Maria (primary), Carlos (secondary)
Response timeline:
14:12 - Alert fires: CheckoutErrorRate 100%
14:13 - Maria receives page (PagerDuty)
14:14 - Maria opens Slack #incident-checkout
14:15 - Maria verifies: 500 errors on all requests
14:16 - Declares Sev-1, opens bridge (Zoom)
14:17 - Invites Carlos (secondary), team lead, DBA
14:18 - Maria investigates: recent deploy?
kubectl rollout history deploy/checkout
-> Deploy v2.4 8 min ago
14:20 - Carlos checks DB: connections saturated
-> Connection pool exhausted
14:22 - Maria decides rollback to v2.3
14:23 - Rollback executed: kubectl rollout undo
14:26 - Service restored, errors at 0%
14:30 - Maria confirms stability for 5 min
14:35 - Closes bridge, declares resolved
14:36 - Creates ticket for post-mortem (48h)
Roles during incident:
| Role | Person | Responsibility |
|------|--------|----------------|
| Incident Commander | Maria | Coordinate, decide |
| Communications | Carlos | Update stakeholders |
| Subject Matter Expert | DBA | Investigate DB |
| Scribe | Auto bot | Timeline in Slack |
Communications:
14:15 - Slack #status: "Investigating Sev-1 checkout"
14:20 - Slack #status: "Root cause identified: deploy v2.4"
14:22 - Slack #status: "Executing rollback"
14:26 - Slack #status: "Service restored"
14:35 - Slack #status: "Incident resolved, post-mortem pending"
Post-mortem (48h):
- Summary: Checkout down 14 min due to defective deploy
- Impact: $28K lost sales, 8K users affected
- Root cause: N+1 query introduced in v2.4, not caught by tests
- 5 Whys:
1. Why did it go down? Connection pool exhausted
2. Why exhausted? N+1 query opened 1000 connections
3. Why not detected? Tests did not cover concurrency
4. Why no tests? No DB integration tests existed
5. Why? CI did not require integration tests
- Actions:
1. Add DB integration test in CI (owner: team, 1 week)
2. Add N+1 detection in CI (owner: platform, 2 weeks)
3. Lower max pool connections (owner: SRE, 3 days)
4. Add pool saturation alert (owner: SRE, 3 days)
Lessons:
- The Incident Commander coordinates, does not investigate
- Communicate early and often
- Rollback is the first option, not the last
- Blameless post-mortem: fix the system, not the blame
- Every action item has an owner and date
How do I prepare a new team for on-call?
Start with shadowing: the new engineer shadows the on-call for 2 weeks without responding to pages. Then they respond to low-severity pages with the senior as backup. After 1 month, they take full rotations with the senior available. Provide a runbook per service. Run game days in staging to practice incident response.
Related Resources
Docker for Developers — A Complete Guide
Learn Docker from the ground up: images, containers, Dockerfiles, networks, volumes, and Docker Compose for local development.
GuideWeb Application Security (OWASP Top 10)
A developer-focused guide to the OWASP Top 10: injection, broken access control, XSS, insecure design, and how to prevent each vulnerability with code examples.
GuideTechnical Documentation Strategy: Docs as Code
A practical guide to treating documentation as code: versioning, review workflows, structure, and tools that keep docs accurate, discoverable, and maintainable.
DocBug Report Template
A structured bug report template to help teams reproduce, triage, and resolve defects faster with clear reproduction steps and expected behavior.
DocService Level Objective (SLO) Document Template
An SLO document template that defines reliability targets, error budgets, and escalation policies for services and platforms.
GuideMonitoring and Alerting — Metrics, Logs, and Dashboards
A practical guide to observability: the three pillars (metrics, logs, traces), RED and USE methods, alert design, and building dashboards that actually help.