Linux SRE · Lesson 4 · 40 min

Incidents, on-call & blameless postmortems

Your phone buzzes: SiteDown. Every SRE, sysadmin and DevOps engineer gets this moment eventually, and what you do in the next ten minutes matters more than any tool. This lesson covers how teams handle incidents calmly: who does what, how to communicate, why you roll back first and debug later, and how to write a blameless postmortem so the same outage never happens twice. Then you'll handle a real one.

You will learn

  • The incident lifecycle: detect → respond → mitigate → resolve → learn
  • Severity levels and incident roles (commander, communications, scribe)
  • How healthy on-call works: rotations, escalation, runbooks
  • Finding “what changed?” fast with journalctl, sudo's log and diff
  • Writing a blameless postmortem, with a timeline and action items

What counts as an incident?

An incident is any unplanned problem that hurts (or is about to hurt) users and needs a coordinated response. Teams sort them by severity, so everyone knows how loud to be:

LevelTypical meaningResponse
SEV1Everything or something critical is down for many users. Data loss or a security breach.Page now, all hands, updates every 15–30 minutes, leadership informed
SEV2A major feature is broken, or it's badly slow for many usersPage the on-call engineer, regular updates
SEV3Minor or partial problem, with a workaroundFix in working hours, as a ticket

Every company sets its own definitions. What matters is that they're written down before the incident, so nobody argues about them in the middle of one.

The lifecycle

  1. Detect: monitoring pages you (ideally before users notice). Time to detect is MTTD.
  2. Respond: acknowledge the page so others know someone has it, open an incident channel, and assess the severity.
  3. Mitigate: stop the pain. Roll back the last change, fail over, or turn off the broken feature. You don't need to understand the bug to undo a change.
  4. Resolve: confirm users are OK, then say so publicly. Time to restore is MTTR.
  5. Learn: write the postmortem, fix the root cause, and track the action items until they're done.
“What changed?”

Most outages are caused by a change: a deploy, a config edit, an upgrade, a certificate expiring. So the fastest first question is what changed recently? Check the deploy log, git log, file times (ls -lt), sudo's log (journalctl -t sudo), and diff against the last known-good copy.

Roles: nobody does everything

For big incidents, one person fixing, coordinating and answering “is it fixed yet?” messages quickly gets overwhelmed. Teams split the work. (The idea comes from emergency services' Incident Command System.)

RoleDoes
Incident Commander (IC)Coordinates and makes decisions, and doesn't type fixes. Keeps asking “what's our next step, and who's doing it?”
Operations / tech leadInvestigates and applies the fix
Communications leadUpdates the status page, customers and leadership on a regular schedule
ScribeWrites down the timeline as it happens (times, findings, actions). The postmortem is built from it.

Small incidents, like the one you're about to handle, are usually one person doing all four. Keep a timeline anyway: date +%H:%M and a notes file are enough.

Good update template

Time (with time zone) · What users see · What we know · What we're doing · Next update at. For example: “14:12 UTC. The ticket site is down for everyone. The web server fails to start after a config change. Rolling that change back now. Next update 14:30.” Updates like “still investigating” are fine. Silence isn't.

On-call that doesn't burn people out

Blameless postmortems

After the incident, write down what happened and what will change. The golden rule is that it's blameless: it asks what allowed this to happen?, never who messed up?

A postmortem's sections
Summary · Impact · Timeline (UTC) · Root cause · What went well
What went wrong · Where we got lucky · Action items (owner, due date)

Tools you'll use in an incident

Rocky / RHEL
systemctl status httpd
sudo apachectl configtest
journalctl -u httpd --since "1 hour ago"
sudo journalctl -t sudo --since "1 hour ago"
sudo grep COMMAND /var/log/secure
ls -lt /etc/httpd/conf.d
diff good.conf current.conf
Ubuntu / Debian
systemctl status apache2
sudo apache2ctl configtest
journalctl -u apache2 --since "1 hour ago"
sudo journalctl -t sudo --since "1 hour ago"
sudo grep COMMAND /var/log/auth.log
ls -lt /etc/apache2/sites-available
diff good.conf current.conf

Practice: you're on call 🚨

The pager just went off. Everything you need is in ~/incident. Follow the runbook, get the site back, keep people updated, then write the postmortem. The timeline in the log says who made the change, but remember: blameless.

Quick check

1. The site broke right after a deploy. You don't know why yet. What's the best first move?

2. Which root cause is written the blameless way?

3. What does the Incident Commander mostly do?

4. Which is a good postmortem action item?

Finished the missions and the quiz? Mark it done to track your progress.