Linux SRE · Lesson 6 · 35 min

Capacity planning & load testing

Every server has a breaking point. Capacity planning means knowing where yours is before your users find it, on the busiest day of the year, at the worst possible moment. You'll load-test a web server to find its limit, learn why it breaks the way it does, and do the simple math that tells you how many months you have before you need a bigger server (or more of them).

You will learn

  • Load testing with ab (and what wrk, hey, k6, JMeter and Locust are for)
  • Throughput vs latency, the knee, and Little's Law
  • Load, stress, soak and spike tests
  • Headroom, target utilization and N+1
  • Forecasting traffic growth and disk fill-up with awk
  • Scaling up vs scaling out
Only test what you own

A load test sends thousands of requests on purpose. Aimed at a website you don't own, it's a denial-of-service attack: against the rules everywhere and against the law in most countries. Test your own servers, preferably a staging copy, and warn your team first. Cloud providers have rules about load testing too, so read them before you start.

Throughput, latency and the knee

Two numbers describe how a server is doing under load:

As you add load, throughput rises and latency stays flat, until some resource (CPU, disk, a lock, a worker pool) is 100% busy. That's the knee. After it, throughput stops growing, and every extra request just waits in a queue, so latency climbs. Push harder still and queues overflow: requests fail.

 req/s                                    latency
  500 ┤      ●──●──●──●──●──●──◐            │                         ●
      │    ●                     ◐ errors    │                     ●
  250 ┤  ●                                   │               ●
      │●                                     │ ●──●──●──●──●
    0 ┼──────────────────────────             ┼──────────────────────────
       1  5  10 50 100 200 400 1000 → users    1  5  10 50 100 200 400 → users
            ↑ the knee                                  ↑ the knee

Little's Law

A simple formula that's true for any system: requests in flight = throughput × latency (L = λ × W). At 500 req/s and 0.1 s each, about 50 requests are in progress at any moment. Flip it around: if the server tops out at 500 req/s and 200 users are clicking at once, each one waits about 200 ÷ 500 = 0.4 s. You'll see exactly that in ab's output.

Kinds of load test

TestQuestion it answers
Load testCan we handle the traffic we expect (for example, peak × 2) and still meet our SLOs?
Stress testWhere's the breaking point, and does it break gracefully (slow down, or return errors) or badly (crash, lose data)?
Soak testDoes it survive hours of steady load? This finds memory leaks, full disks and slowly growing queues.
Spike testWhat happens when traffic jumps 10× in a minute, like a sale or a viral post?

Load testing with ab

ab (ApacheBench) comes with Apache's tools. It's simple and already installed on most web servers:

Rocky / RHEL
sudo dnf install httpd-tools     # has ab (installed with httpd)
Ubuntu / Debian
sudo apt install apache2-utils   # has ab (installed with apache2)
ab -n 1000 -c 10 http://localhost/      # 1000 requests, 10 at a time. Note the / at the end!
ab -q -n 2000 -c 100 http://localhost/ | grep -E 'Requests per second|Failed|  95%'

In the output, look at Requests per second (throughput), Failed requests, and the percentage table (95% of requests finished within that many ms).

ab is fine for a first look, but it hits one URL, from one machine, as fast as it can. For realistic tests, people use wrk or hey (faster), k6 (tests written in JavaScript), Locust (Python), or JMeter. These can simulate whole user journeys, like log in, search, then buy. And the load generator should run on a different machine than the server, or they fight over the same CPU.

Headroom and N+1

Never plan to run at 100%. When a resource is nearly full, latency explodes (you'll see it in the practice). Teams pick a target utilization, often 60–70%, and the difference between current peak and capacity is your headroom.

N+1: if you need N servers to handle the peak, run N+1, so one can fail (or be patched and rebooted) without an outage. With 2 servers behind a load balancer, each must be able to take all the traffic alone, so each should run below 50% at peak.

Forecasting: when do we run out?

Traffic usually grows by a percentage each month (compound growth). Disk usage often grows by a steady amount each day (linear growth). The math is short:

# traffic: grows 8% a month from 178 req/s. When does it reach 350 (70% of 500)?
awk 'BEGIN { print log(350/178) / log(1.08), "months" }'

# disk: average growth per day, from daily samples (date,used_gb,size_gb)
awk -F, 'NR==2 {first=$2} END {print ($2-first)/(NR-2), "GB/day"}' disk-usage.csv

# days until 85% of a 200 GB disk, at 0.35 GB/day, from 133 GB used
awk 'BEGIN { print (200*0.85 - 133) / 0.35, "days" }'

In Prometheus, predict_linear(node_filesystem_avail_bytes[7d], 30*24*3600) does the disk forecast for you. That's how you get alerted weeks before a disk fills, instead of at 3 am. Watch out for seasonality: a school's ticket system is quiet in August and busy in September. Plan for the busy month.

When you run out: scale up or scale out

Scale up (vertical)Scale out (horizontal)
WhatA bigger server: more CPUs and RAMMore servers behind a load balancer
GoodSimple, no code changesAlmost no upper limit, survives a server failing (N+1), can grow and shrink automatically
BadThere's always a biggest server, usually a reboot to resize, and still one point of failureThe app must be built for it: no local state, shared sessions and files

Often the cheapest capacity is efficiency: a cache in front of the app, a missing database index, gzip, a CDN for images. Measure first (lesson 1 of this path!) before buying hardware.

Practice: how much room is left? 📐

The ticket site runs on this 4-CPU server. Find its limit with ab, compare it with real traffic in ~/capacity/traffic.txt, and forecast when it (and the data disk) will run out.

Quick check

1. Going from 50 to 200 concurrent users, requests per second stays at 500, but p95 latency goes from 160 ms to 640 ms. What's happening?

2. Your app runs on 2 servers behind a load balancer, each at 65% CPU at peak. Is that OK?

3. Which test finds a memory leak that only shows up after 6 hours?

4. Where should the load generator run?

Finished the missions and the quiz? Mark it done to track your progress.

Next up: Ansible in depth · Automation at scale, starting with “Variables, loops & handlers”.