Linux SRE · Lesson 5 · 35 min

Limits: cgroups, systemd & the OOM killer

One program with a memory leak can take down everything else on a server. When Linux runs out of RAM, the kernel's OOM killer picks a process to kill, and it often picks the wrong one: your most important, biggest app instead of the small leaky script that caused the trouble. The fix is to put each service in a box with its own limits. Linux calls these boxes control groups (cgroups), and systemd makes them easy to use.

You will learn

  • What happens when Linux runs out of memory, and how to spot an OOM kill
  • cgroups: how systemd puts every service in its own box
  • systemd-cgtop and systemctl status to see who's using what
  • MemoryMax=, MemoryHigh=, CPUQuota= with systemctl set-property and drop-ins
  • OOMScoreAdjust= to protect what matters, and ulimit / LimitNOFILE=

Out of memory

Linux lets programs ask for more memory than really exists, betting they won't all use it at once. Usually that's a good bet. When it loses (RAM and swap are both full), the kernel has two choices: freeze, or kill something. It kills. The OOM killer gives each process a badness score, based mostly on how much memory it uses, and kills the highest.

reportgen invoked oom-killer: gfp_mask=0x140cca(GFP_HIGHUSER_MOVABLE|__GFP_COMP), order=0
Out of memory: Killed process 1830 (java) total-vm:5140000kB, anon-rss:1740800kB, … oom_score_adj:0

Read it carefully. reportgen asked for memory and triggered the OOM killer, but the kernel killed java (the ticket API) because it was the biggest. The victim isn't the culprit. Find the culprit by looking at who's growing, not who's biggest.

Rocky / RHEL
dmesg | grep -i -E 'oom|killed process'
journalctl -k | grep -i oom
Ubuntu / Debian
sudo dmesg | grep -i -E 'oom|killed process'
journalctl -k | grep -i oom

Ubuntu 22.04 and later also run systemd-oomd, which can kill a whole cgroup early, when memory pressure stays high, before the kernel has to step in. Its kills show up in journalctl -u systemd-oomd.

systemd notices too. The unit's status and journal say A process of this unit has been killed by the OOM killer and Failed with result 'oom-kill'.

cgroups: every service in its own box

A cgroup is a group of processes that the kernel measures and limits together: memory, CPU, disk I/O, number of processes. systemd creates one for every service, so a service and all of its children are counted as one. That's also why systemctl stop can reliably stop everything a service started.

systemd-cgtop -m                                   # like top, but per service, sorted by memory
systemctl status reportgen                         # Memory: line = the whole cgroup
cat /sys/fs/cgroup/system.slice/reportgen.service/memory.current
systemd-cgls                                       # the tree of cgroups and their processes

Modern Rocky (9) and Ubuntu (22.04+) both use cgroup v2, a single tree under /sys/fs/cgroup. You can write numbers into those files by hand, but systemd would forget them on restart. Let systemd manage them.

Setting limits with systemd

SettingWhat it does
MemoryMax=300MHard limit. If the service goes over, the kernel OOM-kills inside that cgroup only. Everyone else is safe.
MemoryHigh=250MSoft limit. Above it, the service is slowed down and its memory reclaimed hard. A warning shot before MemoryMax.
CPUQuota=50%At most half of one CPU. 200% means two full CPUs.
CPUWeight= / IOWeight=Relative priority when things are busy (default 100). Like nice, but for a whole service.
TasksMax=Maximum number of processes and threads. Stops fork bombs.
OOMScoreAdjust=-500Makes the kernel less likely to choose this service in a system-wide OOM (range −1000 to 1000, where −1000 means never).
LimitNOFILE=65536How many files and sockets it may have open at once (the ulimit -n of the service)

There are two ways to set them, and both work the same on Rocky and Ubuntu:

# 1. set-property: applies to the running service immediately AND is saved
sudo systemctl set-property reportgen MemoryMax=300M CPUQuota=50%

# 2. a drop-in file: for anything, including non-cgroup settings like OOMScoreAdjust
sudo systemctl edit tickets-api          # opens an editor; add the lines below
[Service]
OOMScoreAdjust=-500
# (or make /etc/systemd/system/tickets-api.service.d/oom.conf yourself, then daemon-reload)
sudo systemctl restart tickets-api

systemctl show reportgen -p MemoryMax,MemoryCurrent,CPUQuotaPerSecUSec
systemctl cat reportgen                   # shows the unit plus every drop-in

set-property writes its drop-ins into /etc/systemd/system.control/. Never edit the vendor's unit file in /usr/lib/systemd/system: a package update would silently overwrite your change. Drop-ins survive updates.

A limit is a seatbelt, not a repair

With MemoryMax, the leaky service gets killed and restarted over and over, but the rest of the server stays healthy. That buys you time. The real fix is still to find and fix the leak (the action item in the postmortem), and NRestarts in systemctl show is a good thing to put on a dashboard.

ulimit: limits for your shell

ulimit -a        # every limit for this shell
ulimit -n        # open files (a classic cause of "Too many open files")

Shell limits come from /etc/security/limits.conf and only apply to login sessions. Services ignore them: for a service, use LimitNOFILE= in its unit. People get this wrong all the time.

Practice: stop the leak from killing the API 🧯

Users report the ticket API keeps crashing. Find out who's really to blame, put the culprit in a box, and protect the API. You'll use timewarp to see what the next two hours would look like.

Quick check

1. The log says reportgen invoked oom-killer … Killed process 1830 (java). Who caused it?

2. You set MemoryMax=300M on reportgen. What happens when it reaches 300 MB?

3. You added nofile 65536 to /etc/security/limits.conf, but your web service still says “Too many open files”. Why?

4. Why not just edit /usr/lib/systemd/system/nginx.service to add a limit?

Finished the missions and the quiz? Mark it done to track your progress.