StackPractices

Tag: on-call

Browse 9 practical software engineering resources tagged with "on-call". Discover code recipes, design patterns, documentation templates, and in-depth guides to help you build, deploy, and maintain production-ready solutions involving on-call. Each resource is written for engineers who ship real systems, with copy-paste examples and practical trade-offs.

On-Call and Incident Readiness

On-call is the practice of having engineers available to respond to production incidents. Effective on-call requires runbooks, escalation paths, alert triage, and a healthy balance with sustainable schedules.

The resources below cover on-call rotations, incident response, alert fatigue, runbooks, and postmortems. Each guide helps you build on-call practices that protect the service and the team.

Every resource includes clear explanations, copy-paste code, and practical warnings. Use them to make informed decisions, avoid production pitfalls, and speed up your delivery. If you are just getting started, read the beginner-friendly articles first; if you are experienced, jump straight to the advanced patterns and architecture guides. New resources are added regularly, so bookmark this page and check back for the latest patterns.

AI LLM Incident Response Runbook

Operational runbook for LLM production incidents: hallucination events, model outages, cost spikes,...

Escalation Policy Template

A template for defining incident severity levels and on-call escalation paths.

On-Call Handoff Template

A template for transferring operational context between on-call shifts including active incidents,...

On-Call Runbook Template

A template documenting common alerts and step-by-step response procedures for on-call engineers.

Service Ownership Document Template

A template for defining who owns a service, what it does, how to operate it, and where to find...

Alert Runbook Template

A standardized runbook for responding to alerts: triage, diagnosis, mitigation, resolution, and...

On-Call and Incident Response Playbook

A practical playbook for on-call engineers: triage, escalation, communication, and postmortems....

Site Reliability Engineering

A practical guide to SRE: defining SLIs, SLOs, and SLAs, managing error budgets, toil reduction,...

Alert Management: On-Call Alerting That Works

A practical guide to alert management: reducing alert fatigue, defining severity levels, escalation...