Incidents, on-call & blameless postmortems
Your phone buzzes: SiteDown. Every SRE, sysadmin and DevOps engineer gets this moment eventually, and what you do in the next ten minutes matters more than any tool. This lesson covers how teams handle incidents calmly: who does what, how to communicate, why you roll back first and debug later, and how to write a blameless postmortem so the same outage never happens twice. Then you'll handle a real one.
You will learn
- The incident lifecycle: detect → respond → mitigate → resolve → learn
- Severity levels and incident roles (commander, communications, scribe)
- How healthy on-call works: rotations, escalation, runbooks
- Finding “what changed?” fast with
journalctl, sudo's log anddiff - Writing a blameless postmortem, with a timeline and action items
What counts as an incident?
An incident is any unplanned problem that hurts (or is about to hurt) users and needs a coordinated response. Teams sort them by severity, so everyone knows how loud to be:
| Level | Typical meaning | Response |
|---|---|---|
| SEV1 | Everything or something critical is down for many users. Data loss or a security breach. | Page now, all hands, updates every 15–30 minutes, leadership informed |
| SEV2 | A major feature is broken, or it's badly slow for many users | Page the on-call engineer, regular updates |
| SEV3 | Minor or partial problem, with a workaround | Fix in working hours, as a ticket |
Every company sets its own definitions. What matters is that they're written down before the incident, so nobody argues about them in the middle of one.
The lifecycle
- Detect: monitoring pages you (ideally before users notice). Time to detect is MTTD.
- Respond: acknowledge the page so others know someone has it, open an incident channel, and assess the severity.
- Mitigate: stop the pain. Roll back the last change, fail over, or turn off the broken feature. You don't need to understand the bug to undo a change.
- Resolve: confirm users are OK, then say so publicly. Time to restore is MTTR.
- Learn: write the postmortem, fix the root cause, and track the action items until they're done.
Most outages are caused by a change: a deploy, a config edit, an upgrade, a certificate expiring. So the fastest first question is what changed recently? Check the deploy log, git log, file times (ls -lt), sudo's log (journalctl -t sudo), and diff against the last known-good copy.
Roles: nobody does everything
For big incidents, one person fixing, coordinating and answering “is it fixed yet?” messages quickly gets overwhelmed. Teams split the work. (The idea comes from emergency services' Incident Command System.)
| Role | Does |
|---|---|
| Incident Commander (IC) | Coordinates and makes decisions, and doesn't type fixes. Keeps asking “what's our next step, and who's doing it?” |
| Operations / tech lead | Investigates and applies the fix |
| Communications lead | Updates the status page, customers and leadership on a regular schedule |
| Scribe | Writes down the timeline as it happens (times, findings, actions). The postmortem is built from it. |
Small incidents, like the one you're about to handle, are usually one person doing all four. Keep a timeline anyway: date +%H:%M and a notes file are enough.
Time (with time zone) · What users see · What we know · What we're doing · Next update at. For example: “14:12 UTC. The ticket site is down for everyone. The web server fails to start after a config change. Rolling that change back now. Next update 14:30.” Updates like “still investigating” are fine. Silence isn't.
On-call that doesn't burn people out
- Rotation: people take turns, often a week each, with a proper handoff (“here's what's flaky right now”). Teams across time zones do follow-the-sun, so nobody gets paged at 3 am.
- Escalation: if the primary doesn't acknowledge within a few minutes, the tool (PagerDuty, Opsgenie, Grafana OnCall…) pages the secondary, then a manager.
- Runbooks: every alert links to a page that says what it means and what to try first. You'll use one in the practice.
- Keep it humane: count the pages. If on-call keeps getting woken up, fixing that becomes the team's top reliability project.
Blameless postmortems
After the incident, write down what happened and what will change. The golden rule is that it's blameless: it asks what allowed this to happen?, never who messed up?
- If people get blamed, they hide mistakes, and the next outage takes longer to understand.
- “Human error” is never a root cause. A person made a typo? Everyone makes typos. The real question is why a typo could take the site down with nothing to catch it.
- Keep asking “why?” (the 5 whys) until you reach something you can fix: Site down → Apache didn't start → config typo → nothing tested the config before the restart → the edit was made by hand, directly on production.
- Every postmortem ends with action items, each with an owner and a due date, tracked like any other work. A postmortem without follow-through is just a story.
Summary · Impact · Timeline (UTC) · Root cause · What went well What went wrong · Where we got lucky · Action items (owner, due date)
Tools you'll use in an incident
systemctl status httpd sudo apachectl configtest journalctl -u httpd --since "1 hour ago" sudo journalctl -t sudo --since "1 hour ago" sudo grep COMMAND /var/log/secure ls -lt /etc/httpd/conf.d diff good.conf current.conf
systemctl status apache2 sudo apache2ctl configtest journalctl -u apache2 --since "1 hour ago" sudo journalctl -t sudo --since "1 hour ago" sudo grep COMMAND /var/log/auth.log ls -lt /etc/apache2/sites-available diff good.conf current.conf
Practice: you're on call 🚨
The pager just went off. Everything you need is in ~/incident. Follow the runbook, get the site back, keep people updated, then write the postmortem. The timeline in the log says who made the change, but remember: blameless.
Quick check
1. The site broke right after a deploy. You don't know why yet. What's the best first move?
✓ Mitigate first. You can debug a rolled-back change at your leisure.
2. Which root cause is written the blameless way?
✓ It points at the process gap you can fix (a configtest step, changes via git and review), not at a person.
3. What does the Incident Commander mostly do?
✓ Keeping hands off the keyboard lets the IC keep the big picture.
4. Which is a good postmortem action item?
✓ Specific, owned and dated, and it makes the mistake impossible instead of asking people to be perfect.