Limits: cgroups, systemd & the OOM killer
One program with a memory leak can take down everything else on a server. When Linux runs out of RAM, the kernel's OOM killer picks a process to kill, and it often picks the wrong one: your most important, biggest app instead of the small leaky script that caused the trouble. The fix is to put each service in a box with its own limits. Linux calls these boxes control groups (cgroups), and systemd makes them easy to use.
You will learn
- What happens when Linux runs out of memory, and how to spot an OOM kill
- cgroups: how systemd puts every service in its own box
systemd-cgtopandsystemctl statusto see who's using whatMemoryMax=,MemoryHigh=,CPUQuota=withsystemctl set-propertyand drop-insOOMScoreAdjust=to protect what matters, andulimit/LimitNOFILE=
Out of memory
Linux lets programs ask for more memory than really exists, betting they won't all use it at once. Usually that's a good bet. When it loses (RAM and swap are both full), the kernel has two choices: freeze, or kill something. It kills. The OOM killer gives each process a badness score, based mostly on how much memory it uses, and kills the highest.
reportgen invoked oom-killer: gfp_mask=0x140cca(GFP_HIGHUSER_MOVABLE|__GFP_COMP), order=0 Out of memory: Killed process 1830 (java) total-vm:5140000kB, anon-rss:1740800kB, … oom_score_adj:0
Read it carefully. reportgen asked for memory and triggered the OOM killer, but the kernel killed java (the ticket API) because it was the biggest. The victim isn't the culprit. Find the culprit by looking at who's growing, not who's biggest.
dmesg | grep -i -E 'oom|killed process' journalctl -k | grep -i oom
sudo dmesg | grep -i -E 'oom|killed process' journalctl -k | grep -i oom
Ubuntu 22.04 and later also run systemd-oomd, which can kill a whole cgroup early, when memory pressure stays high, before the kernel has to step in. Its kills show up in journalctl -u systemd-oomd.
systemd notices too. The unit's status and journal say A process of this unit has been killed by the OOM killer and Failed with result 'oom-kill'.
cgroups: every service in its own box
A cgroup is a group of processes that the kernel measures and limits together: memory, CPU, disk I/O, number of processes. systemd creates one for every service, so a service and all of its children are counted as one. That's also why systemctl stop can reliably stop everything a service started.
systemd-cgtop -m # like top, but per service, sorted by memory systemctl status reportgen # Memory: line = the whole cgroup cat /sys/fs/cgroup/system.slice/reportgen.service/memory.current systemd-cgls # the tree of cgroups and their processes
Modern Rocky (9) and Ubuntu (22.04+) both use cgroup v2, a single tree under /sys/fs/cgroup. You can write numbers into those files by hand, but systemd would forget them on restart. Let systemd manage them.
Setting limits with systemd
| Setting | What it does |
|---|---|
MemoryMax=300M | Hard limit. If the service goes over, the kernel OOM-kills inside that cgroup only. Everyone else is safe. |
MemoryHigh=250M | Soft limit. Above it, the service is slowed down and its memory reclaimed hard. A warning shot before MemoryMax. |
CPUQuota=50% | At most half of one CPU. 200% means two full CPUs. |
CPUWeight= / IOWeight= | Relative priority when things are busy (default 100). Like nice, but for a whole service. |
TasksMax= | Maximum number of processes and threads. Stops fork bombs. |
OOMScoreAdjust=-500 | Makes the kernel less likely to choose this service in a system-wide OOM (range −1000 to 1000, where −1000 means never). |
LimitNOFILE=65536 | How many files and sockets it may have open at once (the ulimit -n of the service) |
There are two ways to set them, and both work the same on Rocky and Ubuntu:
# 1. set-property: applies to the running service immediately AND is saved sudo systemctl set-property reportgen MemoryMax=300M CPUQuota=50% # 2. a drop-in file: for anything, including non-cgroup settings like OOMScoreAdjust sudo systemctl edit tickets-api # opens an editor; add the lines below [Service] OOMScoreAdjust=-500 # (or make /etc/systemd/system/tickets-api.service.d/oom.conf yourself, then daemon-reload) sudo systemctl restart tickets-api systemctl show reportgen -p MemoryMax,MemoryCurrent,CPUQuotaPerSecUSec systemctl cat reportgen # shows the unit plus every drop-in
set-property writes its drop-ins into /etc/systemd/system.control/. Never edit the vendor's unit file in /usr/lib/systemd/system: a package update would silently overwrite your change. Drop-ins survive updates.
With MemoryMax, the leaky service gets killed and restarted over and over, but the rest of the server stays healthy. That buys you time. The real fix is still to find and fix the leak (the action item in the postmortem), and NRestarts in systemctl show is a good thing to put on a dashboard.
ulimit: limits for your shell
ulimit -a # every limit for this shell ulimit -n # open files (a classic cause of "Too many open files")
Shell limits come from /etc/security/limits.conf and only apply to login sessions. Services ignore them: for a service, use LimitNOFILE= in its unit. People get this wrong all the time.
Practice: stop the leak from killing the API 🧯
Users report the ticket API keeps crashing. Find out who's really to blame, put the culprit in a box, and protect the API. You'll use timewarp to see what the next two hours would look like.
Quick check
1. The log says reportgen invoked oom-killer … Killed process 1830 (java). Who caused it?
✓ The victim isn't the culprit. “Invoked” only means reportgen happened to ask for memory at that moment.
2. You set MemoryMax=300M on reportgen. What happens when it reaches 300 MB?
✓ That's the whole point of a cgroup limit: it contains the damage.
3. You added nofile 65536 to /etc/security/limits.conf, but your web service still says “Too many open files”. Why?
✓ Check what a service really got with systemctl show NAME -p LimitNOFILE.
4. Why not just edit /usr/lib/systemd/system/nginx.service to add a limit?
✓ /usr is the vendor's, /etc is yours. That's the same rule you learned in the filesystem lesson.