Linux SRE: keep it up, find out why
Something is always a little broken in production. Site Reliability Engineering is the craft of knowing how broken, finding out why quickly, and making sure the same thing doesn't bite you twice. You'll measure before you guess, stay calm when the pager goes off, and turn every outage into a lesson. If you liked tracking down the hidden miner or the full disk in Linux Sysadmin, this path is more of that, with better tools.
Is this path for you?
You're ready if these sound familiar (each links to the lesson that teaches it):
- Finding and stopping processes:
ps,top,kill(Linux Sysadmin, lesson 6) - Reading logs with
journalctland/var/log(Linux Sysadmin, lesson 4) - Services and unit files with
systemctl(Linux Basics, lesson 15) - Checking disks and filesystems with
dfanddu(Linux Sysadmin, lesson 7) - Scheduling jobs with cron (Linux Sysadmin, lesson 5)
If some of those are fuzzy, Linux Sysadmin covers all of them. Take your time there first.
What's ahead
| Lesson | You'll learn to… | The challenge |
|---|---|---|
| Measure | ||
| lesson 1: the 60-second checklist | use the USE method with uptime, vmstat, mpstat, pidstat, iostat, free and sar | A slow server. Find the disk hog and the timer that keeps bringing it back. |
| lesson 2: SLIs, SLOs, SLAs & KPIs | pick what to measure, set targets in “nines”, spend error budgets, read percentiles | Grade a day of real traffic against its SLOs, using awk |
| lesson 3: Prometheus | run node_exporter and Prometheus, ask questions in PromQL, write alert rules | Build monitoring from scratch, then make an alert fire |
| Respond & plan | ||
| lesson 4: incidents & postmortems | run an incident, communicate, roll back first, write a blameless postmortem | You're on call and the site is down |
| lesson 5: cgroups & the OOM killer | read OOM kills, box services in with MemoryMax, CPUQuota and OOMScoreAdjust | A leak keeps getting the wrong service killed |
| lesson 6: capacity & load testing | load-test with ab, use Little's Law and N+1, forecast growth | How many months until this server is too small? |
How this path is different
- Measure first. Every lesson starts with numbers, not guesses. “It feels slow” becomes “disk writes take 150 ms and the queue is 150 deep.”
- Mitigate, then fix the cause. Stop the pain for users first, then find the real reason so it never comes back.
- No blame. When something breaks, the question is “what let this happen?”, not “who did it?” People who aren't afraid of blame tell you what really happened.
The best troubleshooters aren't the ones who know every command. They're the ones who stay curious and methodical when everything is on fire. Follow the checklist, write down what you see, and change one thing at a time. That habit will serve you longer than any tool in this path.