Performance triage: the 60-second checklist
“The website is slow.” Every SRE hears it, usually with no other details. You could guess: restart something, add memory, blame the network. Or you could measure. This lesson gives you a checklist of ten commands that, in about a minute, tells you which part of the server is struggling and which process is causing it. Then you'll use it on a server that's slow right now.
You will learn
- What an SRE does, and the USE method for checking any resource
- A 10-command triage checklist:
uptime,dmesg,vmstat,mpstat,pidstat,iostat,free,sar,top - Installing
sysstaton both families - Reading load average,
%iowait,%utilandawaitwithout panicking - Mitigate first, then fix the cause: stop the bleeding, then find why it happened
What's an SRE?
Site Reliability Engineering started at Google in the early 2000s. The idea: run production like an engineering problem. Measure how reliable things are, automate the boring fixes, and when something breaks, find out why so it doesn't happen again. The job title is newer than Linux sysadmin work, but the core skill is the same one you practised in Linux Sysadmin: read the evidence, change one thing, check again.
The USE method
Brendan Gregg's USE method is a way to never forget to check something. For every resource (CPU, memory, disk, network), ask three questions:
| Question | Example | |
|---|---|---|
| Utilization | How busy is it? | The disk was busy 99% of the time |
| Saturation | Is work queuing up, waiting its turn? | 150 requests were waiting in the disk's queue |
| Errors | Is anything failing? | The kernel logged disk errors or blocked tasks |
A resource can be 100% busy and fine, like a CPU crunching a video. The trouble starts with saturation: that's where waiting, and slowness, come from.
The checklist
This list comes from Brendan Gregg's Linux Performance Analysis in 60,000 Milliseconds, written at Netflix. Run it top to bottom, every time, before you guess.
uptime # 1. load averages: is it busy, and getting better or worse? dmesg | tail # 2. kernel errors (sudo on Ubuntu) vmstat 1 5 # 3. run queue, memory, swap, disk I/O, CPU split mpstat -P ALL 1 3 # 4. every CPU: is one hot? is time spent waiting on I/O? pidstat 1 3 # 5. which PROCESSES use CPU iostat -xz 1 3 # 6. which DISKS are busy, and how slow they respond free -m # 7. memory sar -n DEV 1 3 # 8. network throughput sar -n TCP,ETCP 1 3 # 9. TCP connections and retransmits top # 10. the overview, to check your story
The two numbers mean “every 1 second, 3 times”. Without the count, most of these keep going until you press Ctrl+C.
Install sysstat
vmstat, free and top come with every install. mpstat, pidstat, iostat and sar are in the sysstat package, which has the same name on both families:
sudo dnf install sysstat
sudo apt update sudo apt install sysstat
Install it on every server before you need it. In the middle of an incident is a bad time to find out a tool is missing.
Reading the numbers
1. uptime: load average
05:12:17 up 1:12, 1 user, load average: 6.12, 5.84, 5.02
Three averages: the last 1, 5 and 15 minutes. Compare them with the number of CPUs (nproc). Load 6 on 4 CPUs means more work than the CPUs can serve. Reading left to right tells you the trend: here the 1-minute number is highest, so it's getting worse.
On Linux, load counts processes that are running and processes stuck waiting for the disk (state D, “uninterruptible sleep”). High load with idle CPUs usually means a disk problem.
2. dmesg: did the kernel complain?
INFO: task tar:2419 blocked for more than 122 seconds.
Look for the OOM killer (memory ran out), disk or filesystem errors, network card resets, and blocked for more than 120 seconds, which means a process waited on the disk for over two minutes. On Ubuntu, use sudo dmesg.
3. vmstat 1 5: the whole machine, second by second
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
1 0 0 98304 4212 2802584 0 0 1480 2710 240 410 1 1 94 4 0 ← since boot: skip it
1 4 0 94524 4212 2788233 0 0 63319 117420 1781 3733 8 6 24 62 0
| Column | Means | Worry when… |
|---|---|---|
r | processes running or waiting for a CPU | bigger than the CPU count |
b | processes blocked waiting on I/O | more than 0 for a while |
si / so | swap in / out | not 0: you're out of memory |
bi / bo | blocks read from / written to disk | much higher than normal for this server |
us sy id wa st | CPU time: user, system (kernel), idle, waiting for I/O, stolen by the hypervisor | wa high = disk; st high = a noisy neighbour on your cloud host |
4 & 5. mpstat and pidstat: which CPU, which process
mpstat -P ALL 1 shows each CPU. One CPU at 100% and the rest idle means a single-threaded program has hit its limit. pidstat 1 lists only the processes that were busy during each second. pidstat -d 1 does the same for disk: how many kB each process read and wrote per second.
6. iostat -xz 1: the disks
The table is wide. Focus on four columns:
r_await/w_await: milliseconds per read and write, including time spent waiting in the queue. A healthy SSD answers in about 1 ms, and even a slow disk in about 10. Hundreds means requests are stuck in a queue.aqu-sz: the average queue length. Saturation.%util: how much of the time the disk was busy. Close to 100 means it's maxed out. (Modern SSDs and RAID arrays can handle many requests at once, so for them, trustawaitandaqu-szmore than%util.)
7. free -m: memory
total used free shared buff/cache available Mem: 3671 834 96 16 2741 2737
Only 96 MB “free” looks scary, but it isn't. Linux uses spare memory as a disk cache and gives it back the moment a program needs it. The number that matters is available. Worry when available is small and vmstat shows swapping.
8 – 10. sar and top
sar -n DEV 1 shows network traffic per interface, and sar -n TCP,ETCP 1 shows new connections and retransmits (lost packets being sent again). When those are quiet, you've ruled out the network, and that's a useful answer too. Finish with top: in its %Cpu(s) line, wa is I/O wait, and in the S column, D marks processes stuck on the disk.
sar remembers
With an interval, sar shows live numbers. With no interval, it reads the history that sysstat saves every 10 minutes, so you can ask “what was this server doing at 3 am?” Collection has to be switched on: sudo systemctl enable --now sysstat. On Ubuntu, first set ENABLED="true" in /etc/default/sysstat.
Mitigate first, then fix the cause
When users are hurting, first stop the bleeding. Stop the job, move the traffic, roll back the deploy. Then, calmly, find the root cause: why was that job running now? Fixing the symptom and walking away means it will happen again tomorrow. You'll write up what you find in a postmortem, a lesson that's coming later in this path.
ps -o pid,ppid,stat,cmd -p 2419 # who is its parent (PPID)? systemctl status 2419 # which service is this process part of? sudo systemctl stop nightly-backup # stop the bleeding systemctl list-timers # then: why was it running at all? systemctl cat nightly-backup.timer
Stopping a service stops everything it started. systemd keeps each service's processes together in a cgroup, so the tar and gzip it launched stop too. kill on just the parent script would have left tar running.
systemd timers: OnCalendar
Timers are systemd's version of cron. When a timer should fire is set with OnCalendar=:
OnCalendar= | Fires |
|---|---|
hourly | at the start of every hour |
daily | every day at midnight |
*-*-* 02:00:00 | every day at 2 am (year-month-day hour:minute:second) |
Mon *-*-* 09:00:00 | Mondays at 9 am |
On a real server, systemd-analyze calendar '*-*-* 02:00:00' checks an expression and shows when it will fire next. After editing a unit file, run sudo systemctl daemon-reload so systemd reads the new version.
Even at 2 am, a backup can slow down the rest of the server. Adding Nice=19 and IOSchedulingClass=idle to the [Service] section tells Linux to give it CPU time and disk time only when nothing else wants them.
Practice: why is this server slow? 🐢
Users say the website on this server has been slow all day. Run the checklist, find the cause, stop it, and fix it so it doesn't come back.
Quick check
1. A 4-CPU server shows load average 6.1, but top says the CPUs are 22% idle. What's the most likely story?
✓ On Linux, load includes tasks in D state (waiting for I/O). High load with idle CPU usually points at the disk.
2. free -m shows only 96 MB free and 2.7 GB of buff/cache. Should you add memory?
✓ Unused memory is wasted memory, so Linux borrows it as disk cache.
3. Why ignore the first line of vmstat 1 and the first report of iostat -xz 1?
✓ The lines after it are live samples, one per interval.
4. You stopped the backup and the site is fast again. What's left to do?
✓ Mitigation stops the pain. The root-cause fix stops it from coming back, and you still need backups!