Linux Sysadmin · Lesson 6 · 25 min

Processes & disk space

“The server is slow.” “The website crashed, it says the disk is full.” These are the two calls every sysadmin gets. Today you'll learn the tools to answer them, and then handle two realistic incidents.

You will learn

  • What processes are, and how to see them: ps, top, pgrep
  • How to stop them: kill, kill -9, pkill, killall
  • How to check memory and CPU load: free, uptime, nproc
  • How to find what's eating your disk: df, du, find, and lsof for the space that won't come back

Processes: programs that are running

Every running program is a process with a number called a PID. Some belong to you, and many belong to root or to service accounts. Here's how to see them all:

Same on both
ps aux                  # every process: user, PID, %CPU, %MEM, command
ps aux | grep sshd      # find one
pgrep -l sshd           # just PIDs and names
top                     # live dashboard, sorted by CPU (q to quit)
Rocky / RHEL
sudo dnf install epel-release
sudo dnf install htop

The friendlier, colorful htop lives in EPEL.

Ubuntu / Debian
htop

htop comes pre-installed on Ubuntu Server.

How busy is the machine?

Same on both
uptime      # ... load average: 0.08, 0.03, 0.01
nproc       # how many CPU cores
free -h     # memory: look at the "available" column

The load average shows how many things wanted the CPU, averaged over the last 1, 5 and 15 minutes. Compare it to nproc: a load of 4.0 on a 4-core machine means it's completely busy.

Stopping a process

Same on both
kill 1234        # politely: "please stop" (signal 15, TERM)
kill -9 1234     # forcefully: no cleanup (signal 9, KILL). Last resort.
pkill firefox    # by name instead of number (matches part of the name)
killall firefox  # by EXACT name

You can only kill your own processes. Other people's need sudo. For services, don't kill them. Use systemctl stop / restart instead.

Disk space

Same on both
df -h                          # how full is each disk? watch Use%
df -hT                         # ...with the filesystem type
du -sh *                       # size of each thing in this folder
sudo du -h -d 1 /var | sort -h # biggest folders under /var, largest last
sudo find / -size +1G          # any file over 1 GB
Spot the difference

df -T shows the filesystem type: Rocky formats its disks as xfs and Ubuntu as ext4. Rocky also puts the system on LVM volumes named like /dev/mapper/rl-root, and Ubuntu uses ubuntu--vg-ubuntu--lv. For everyday use they behave the same.

The trap: “I deleted it but the disk is still full!”

If a running program still has a file open, rm only removes the name from the folder. The data stays on disk, invisible, until that program closes it. df still says 100% and du can't find it. lsof (“list open files”) can:

Same on both
sudo lsof +L1                        # open files with 0 links, i.e. deleted but still held
sudo lsof /var/log/app/debug.log     # who has this file open?
sudo systemctl restart myapp         # restarting (or killing) the program frees the space
COMMAND  PID USER   FD   TYPE DEVICE    SIZE/OFF NLINK    NODE NAME
myapp   2210 root    3w   REG  253,0 15247133901     0 1048612 /var/log/app/debug.log (deleted)
Empty it instead of deleting it

sudo truncate -s 0 FILE empties a file in place, so the space is freed right away, even while the program keeps writing. Then work out why it grew. Logs that grow forever usually mean something is misconfigured, or that nobody set up log rotation (the logrotate tool, which compresses and deletes old logs on a schedule).

Incident 1: “The server is really slow” 🐌

The school server has felt sluggish since yesterday. Investigate. When you find something suspicious, remember that attackers often set things up to come back, and cron (last lesson) is a favorite way to do it.

In real life

A crypto-miner on your server means someone got in, and in this scenario last shows a login from an address that isn't yours. Killing the process isn't enough. You'd also change passwords, check ~/.ssh/authorized_keys for keys you didn't add, update the system, and tell whoever is responsible for security. On a seriously compromised machine, pros often reinstall from scratch and restore from backups (lesson 8).

Incident 2: “The disk is full” 💾

A fresh, separate server. Programs are crashing with “No space left on device.” Find the space hog and fix it.

Quick check

1. uptime shows a load average of 3.9 and nproc says 4. What does that mean?

2. You kill a suspicious process and it's back 5 minutes later. What should you check?

3. df -h shows / at 98%. Which command finds the biggest folders?

Finished the missions and the quiz? Mark it done to track your progress.