SLIs, SLOs, SLAs & KPIs: measuring reliability
“Is the site reliable?” sounds like a yes-or-no question, but it isn't. Nothing is up 100% of the time, not Google, not your bank, not this website. The real question is how reliable is reliable enough, and how you'd know. SRE answers it with a handful of acronyms that sound scary and are actually simple: measure the right thing (SLI), pick a target (SLO), and spend the gap between them wisely (error budget).
You will learn
- The difference between an SLI, an SLO, an SLA and a KPI
- What “three nines” really means in minutes
- Error budgets and burn rate: how SLOs let teams move fast and stay reliable
- Why percentiles beat averages
- Measuring real SLIs from a web server log with
awk
Four acronyms, one idea
| Stands for | It's… | Example | |
|---|---|---|---|
| SLI | Service Level Indicator | A measurement of how the service behaves, as users see it. Usually “good events ÷ all events”. | 99.45% of requests yesterday returned a non-error status |
| SLO | Service Level Objective | A target for an SLI, over a time window. Set by the team. | 99.5% of requests succeed, measured over 30 days |
| SLA | Service Level Agreement | A promise in a contract, with a penalty (usually money back) if it's broken. Set by lawyers and sales. | “99.9% monthly uptime, or you get 10% credit” |
| KPI | Key Performance Indicator | A business number leaders watch. Not reliability-specific. | daily active users, sign-ups, revenue, tickets closed |
An easy way to remember them: the Indicator is what you measure, the Objective is what you aim for, and the Agreement is what you've promised. SLOs are always stricter than SLAs, so the team gets a warning before a customer gets a refund. KPIs are the business's scoreboard. Reliability affects them (a slow checkout means fewer sales), which is why SREs care about them too.
MTTD (mean time to detect), MTTR (mean time to restore), change failure rate (how many deploys cause a problem) and deployment frequency. The last two, with lead time and MTTR, are the four DORA metrics used to measure DevOps teams.
Choosing good SLIs
A good SLI measures what a user feels, not what a server feels. “CPU at 40%” isn't an SLI: no user has ever noticed your CPU. “Checkout pages load in under a second” is. Google's SRE book suggests starting from the four golden signals:
| Signal | Question | Example SLI |
|---|---|---|
| Latency | How long do requests take? | % of requests served in ≤ 300 ms |
| Traffic | How much demand is there? | requests per second |
| Errors | How many requests fail? | % of requests with status ≥ 500 |
| Saturation | How full is the system? | disk 85% used, queue 150 deep |
Traffic and saturation help you explain problems. Latency and errors are the ones users feel, so they make the best SLOs. Note that a 404 Not Found is usually the user's mistake (a bad link), so availability SLIs count only 5xx responses as failures.
How many nines?
Availability is usually written as “nines”. Each extra nine allows ten times less downtime, and it usually costs a lot more to reach:
| Target | Name | Allowed downtime per 30 days |
|---|---|---|
| 99% | two nines | 7 hours 12 minutes |
| 99.5% | two and a half | 3 hours 36 minutes |
| 99.9% | three nines | 43 minutes 12 seconds |
| 99.99% | four nines | 4 minutes 19 seconds |
| 99.999% | five nines | 26 seconds |
The math is one line: 30 days × 24 × 60 × (1 − 0.999) = 43.2 minutes. Four nines leaves no time for a human to even wake up, so it needs automatic failover. Don't aim higher than your users need: if their home Wi-Fi is only 99% reliable, they'll never notice the difference between your 99.99% and 99.999%.
Error budgets
Here's the clever part. If the SLO is 99.5%, then 0.5% of requests are allowed to fail. That 0.5% is the error budget, and the team gets to spend it:
- Budget left? Ship features, deploy on Fridays, try risky upgrades. Some will cause small failures, and that's fine: it's what the budget is for.
- Budget spent? Slow down. Freeze risky changes and work on reliability until the budget refills.
This ends the classic fight between developers (“ship faster!”) and operations (“change nothing!”). The data decides, and everyone agreed on the rules ahead of time, in an error budget policy.
Burn rate
Burn rate is how fast you're spending the budget. A burn rate of 1 uses exactly the whole budget in 30 days. A burn rate of 10 would use it all in 3 days. Good alerts fire on fast burn (page someone now) and slow burn (open a ticket), instead of on every single error.
Averages lie, percentiles don't
Suppose 99 requests take 50 ms and one takes 10 seconds. The average is about 150 ms, which describes nobody's actual experience. Percentiles do better: sort all the times, then pick a position.
- p50 (the median): half the requests were faster than this. The “typical” user.
- p95: 95% were faster. 1 in 20 users waits longer.
- p99: 1 in 100 waits longer. On a busy site that's thousands of real people every day, and often your most active users, who make the most requests.
In the practice log, the average is 109 ms and p95 is 195 ms, both comfortably fast. But p99 is 870 ms, and that's where the day's incident hides.
Measuring SLIs from a log
Before you have fancy monitoring (next lesson), your web server's access log already holds everything you need. In the practice log, each line ends with the response time in milliseconds:
198.51.100.8 - - [27/Sep/2026:00:00:26 +0000] "GET / HTTP/1.1" 200 2047 "-" "Mozilla/5.0" 88
$1 = client IP $4 = time $7 = path $9 = status $NF = the last field (ms)
This is the same combined log format as on your web servers. The response time is an extra you switch on: %D (microseconds) in Apache's LogFormat, or $request_time (seconds) in nginx's log_format.
# availability SLI: % of requests that were NOT a 5xx awk '$9 < 500 {good++} END {printf "%.2f%%\n", 100*good/NR}' access.log # latency SLI: % of requests served in 300 ms or less awk '$NF <= 300 {fast++} END {printf "%.2f%%\n", 100*fast/NR}' access.log # percentiles: sort the times, then pick positions awk '{print $NF}' access.log | sort -n | awk '{t[NR]=$1} END {print "p50", t[int(NR*0.5)], "p95", t[int(NR*0.95)], "p99", t[int(NR*0.99)]}' # when did the errors happen? (characters 14-18 of $4 are HH:MM) awk '$9 >= 500 {print substr($4, 14, 5)}' access.log | uniq -c
In awk, NR is the number of lines read so far, so in the END block it's the total. good++ counts, and printf formats the result to two decimal places.
gawk on Rocky, mawk on Ubuntu
Rocky's awk is GNU awk, and Ubuntu's is mawk (smaller and faster). Everything in this lesson works the same in both. They only differ in extras, like gawk's asort().
Practice: grade yesterday 📏
The ticket app's SLOs are in ~/slo.txt, and yesterday's log is in ~/logs/access.log. Did the team meet its targets? What happened, and how much budget did it cost?
Quick check
1. “99.9% of checkout requests succeed over 30 days” is…
✓ The SLI is the measured success rate. Adding a target and a window makes it an objective.
2. Your SLA promises customers 99.9%. What should your internal SLO be?
✓ The SLO is your early warning. The SLA is the line you must never cross.
3. The error budget has plenty left and the team wants to ship a risky new feature. What does an error budget policy say?
✓ And if the budget runs out, the same policy says: slow down and fix reliability first.
4. Average latency is 109 ms, but p99 is 870 ms. What does that tell you?
✓ Look at the tail. That's where the unhappy users (and the incidents) are.