Final challenge: on call
You're on call for the tickets service. The pager just went off: the site is down. While you're in there, people have also been complaining that the server is sluggish, and the ticket API keeps dying at random. Deal with the outage the way an SRE team would, then fix the other two for good.
No step-by-step instructions: the objectives say what must be true when you're done. Start with incident/page.txt and the runbook. Communicate while you work: users read incident/updates.txt. The cheat sheet and search (/) are allowed, and hints are below.
The brief
- The site is back up. The web server is running again with a working configuration.
- Users were kept informed.
incident/updates.txthas at least two updates from you, and the last one says RESOLVED. - The server is fast again, and the job that slowed it down now runs at 02:00, when nobody's using the site.
- The API stops dying. The service that leaks memory is capped, and the ticket API is protected from the out-of-memory killer.
- Learned from it.
incident/postmortem.mdis filled in (no TODOs left), with a timeline of at least three times and at least one action item.
Stuck? Hints
Open only as many as you need.
1. The site
The runbook's first step is a config test. Something was edited by hand an hour ago: journalctl -t sudo shows who and what, and there's a .bak of the file. Roll back first, investigate later.
Still stuck: lesson 4.
2. Updates
Append lines with echo "$(date +%H:%M) …" >> incident/updates.txt: one when you know what's wrong, one saying RESOLVED when it's fixed.
3. Fast again
vmstat, iostat and pidstat -d point at a process; systemctl status PID says which unit started it. Stop it, then change the timer's OnCalendar and daemon-reload.
Still stuck: lesson 1.
4. The API
The kernel log (dmesg) says who the OOM killer picked, and systemd-cgtop -m shows who's growing. systemctl set-property … MemoryMax= puts a service in a box; a drop-in with OOMScoreAdjust=-500 protects another.
Still stuck: lesson 5.
5. Postmortem
Copy incident/postmortem-template.md to incident/postmortem.md and replace every TODO: summary, impact, a timeline with times (05:39), the root cause, and action items as - [ ] …. Blameless: what failed, not who.