Linux SRE · Lesson 1 · 30 min

Performance triage: the 60-second checklist

“The website is slow.” Every SRE hears it, usually with no other details. You could guess: restart something, add memory, blame the network. Or you could measure. This lesson gives you a checklist of ten commands that, in about a minute, tells you which part of the server is struggling and which process is causing it. Then you'll use it on a server that's slow right now.

You will learn

  • What an SRE does, and the USE method for checking any resource
  • A 10-command triage checklist: uptime, dmesg, vmstat, mpstat, pidstat, iostat, free, sar, top
  • Installing sysstat on both families
  • Reading load average, %iowait, %util and await without panicking
  • Mitigate first, then fix the cause: stop the bleeding, then find why it happened

What's an SRE?

Site Reliability Engineering started at Google in the early 2000s. The idea: run production like an engineering problem. Measure how reliable things are, automate the boring fixes, and when something breaks, find out why so it doesn't happen again. The job title is newer than Linux sysadmin work, but the core skill is the same one you practised in Linux Sysadmin: read the evidence, change one thing, check again.

The USE method

Brendan Gregg's USE method is a way to never forget to check something. For every resource (CPU, memory, disk, network), ask three questions:

QuestionExample
UtilizationHow busy is it?The disk was busy 99% of the time
SaturationIs work queuing up, waiting its turn?150 requests were waiting in the disk's queue
ErrorsIs anything failing?The kernel logged disk errors or blocked tasks

A resource can be 100% busy and fine, like a CPU crunching a video. The trouble starts with saturation: that's where waiting, and slowness, come from.

The checklist

This list comes from Brendan Gregg's Linux Performance Analysis in 60,000 Milliseconds, written at Netflix. Run it top to bottom, every time, before you guess.

Both families
uptime                 # 1. load averages: is it busy, and getting better or worse?
dmesg | tail           # 2. kernel errors (sudo on Ubuntu)
vmstat 1 5             # 3. run queue, memory, swap, disk I/O, CPU split
mpstat -P ALL 1 3      # 4. every CPU: is one hot? is time spent waiting on I/O?
pidstat 1 3            # 5. which PROCESSES use CPU
iostat -xz 1 3         # 6. which DISKS are busy, and how slow they respond
free -m                # 7. memory
sar -n DEV 1 3         # 8. network throughput
sar -n TCP,ETCP 1 3    # 9. TCP connections and retransmits
top                    # 10. the overview, to check your story

The two numbers mean “every 1 second, 3 times”. Without the count, most of these keep going until you press Ctrl+C.

Install sysstat

vmstat, free and top come with every install. mpstat, pidstat, iostat and sar are in the sysstat package, which has the same name on both families:

Rocky / RHEL
sudo dnf install sysstat
Ubuntu / Debian
sudo apt update
sudo apt install sysstat

Install it on every server before you need it. In the middle of an incident is a bad time to find out a tool is missing.

Reading the numbers

1. uptime: load average

 05:12:17 up 1:12,  1 user,  load average: 6.12, 5.84, 5.02

Three averages: the last 1, 5 and 15 minutes. Compare them with the number of CPUs (nproc). Load 6 on 4 CPUs means more work than the CPUs can serve. Reading left to right tells you the trend: here the 1-minute number is highest, so it's getting worse.

Linux load is not just CPU

On Linux, load counts processes that are running and processes stuck waiting for the disk (state D, “uninterruptible sleep”). High load with idle CPUs usually means a disk problem.

2. dmesg: did the kernel complain?

INFO: task tar:2419 blocked for more than 122 seconds.

Look for the OOM killer (memory ran out), disk or filesystem errors, network card resets, and blocked for more than 120 seconds, which means a process waited on the disk for over two minutes. On Ubuntu, use sudo dmesg.

3. vmstat 1 5: the whole machine, second by second

procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 1  0      0  98304   4212 2802584    0    0  1480  2710  240  410  1  1 94  4  0   ← since boot: skip it
 1  4      0  94524   4212 2788233    0    0 63319 117420 1781 3733  8  6 24 62  0
ColumnMeansWorry when…
rprocesses running or waiting for a CPUbigger than the CPU count
bprocesses blocked waiting on I/Omore than 0 for a while
si / soswap in / outnot 0: you're out of memory
bi / boblocks read from / written to diskmuch higher than normal for this server
us sy id wa stCPU time: user, system (kernel), idle, waiting for I/O, stolen by the hypervisorwa high = disk; st high = a noisy neighbour on your cloud host

4 & 5. mpstat and pidstat: which CPU, which process

mpstat -P ALL 1 shows each CPU. One CPU at 100% and the rest idle means a single-threaded program has hit its limit. pidstat 1 lists only the processes that were busy during each second. pidstat -d 1 does the same for disk: how many kB each process read and wrote per second.

6. iostat -xz 1: the disks

The table is wide. Focus on four columns:

7. free -m: memory

               total        used        free      shared  buff/cache   available
Mem:            3671         834          96          16        2741        2737

Only 96 MB “free” looks scary, but it isn't. Linux uses spare memory as a disk cache and gives it back the moment a program needs it. The number that matters is available. Worry when available is small and vmstat shows swapping.

8 – 10. sar and top

sar -n DEV 1 shows network traffic per interface, and sar -n TCP,ETCP 1 shows new connections and retransmits (lost packets being sent again). When those are quiet, you've ruled out the network, and that's a useful answer too. Finish with top: in its %Cpu(s) line, wa is I/O wait, and in the S column, D marks processes stuck on the disk.

sar remembers

With an interval, sar shows live numbers. With no interval, it reads the history that sysstat saves every 10 minutes, so you can ask “what was this server doing at 3 am?” Collection has to be switched on: sudo systemctl enable --now sysstat. On Ubuntu, first set ENABLED="true" in /etc/default/sysstat.

Mitigate first, then fix the cause

When users are hurting, first stop the bleeding. Stop the job, move the traffic, roll back the deploy. Then, calmly, find the root cause: why was that job running now? Fixing the symptom and walking away means it will happen again tomorrow. You'll write up what you find in a postmortem, a lesson that's coming later in this path.

Both families
ps -o pid,ppid,stat,cmd -p 2419      # who is its parent (PPID)?
systemctl status 2419                # which service is this process part of?
sudo systemctl stop nightly-backup   # stop the bleeding
systemctl list-timers                # then: why was it running at all?
systemctl cat nightly-backup.timer

Stopping a service stops everything it started. systemd keeps each service's processes together in a cgroup, so the tar and gzip it launched stop too. kill on just the parent script would have left tar running.

systemd timers: OnCalendar

Timers are systemd's version of cron. When a timer should fire is set with OnCalendar=:

OnCalendar=Fires
hourlyat the start of every hour
dailyevery day at midnight
*-*-* 02:00:00every day at 2 am (year-month-day hour:minute:second)
Mon *-*-* 09:00:00Mondays at 9 am

On a real server, systemd-analyze calendar '*-*-* 02:00:00' checks an expression and shows when it will fire next. After editing a unit file, run sudo systemctl daemon-reload so systemd reads the new version.

Make heavy jobs polite

Even at 2 am, a backup can slow down the rest of the server. Adding Nice=19 and IOSchedulingClass=idle to the [Service] section tells Linux to give it CPU time and disk time only when nothing else wants them.

Practice: why is this server slow? 🐢

Users say the website on this server has been slow all day. Run the checklist, find the cause, stop it, and fix it so it doesn't come back.

Quick check

1. A 4-CPU server shows load average 6.1, but top says the CPUs are 22% idle. What's the most likely story?

2. free -m shows only 96 MB free and 2.7 GB of buff/cache. Should you add memory?

3. Why ignore the first line of vmstat 1 and the first report of iostat -xz 1?

4. You stopped the backup and the site is fast again. What's left to do?

Finished the missions and the quiz? Mark it done to track your progress.