Linux SRE · Lesson 2 · 35 min

SLIs, SLOs, SLAs & KPIs: measuring reliability

“Is the site reliable?” sounds like a yes-or-no question, but it isn't. Nothing is up 100% of the time, not Google, not your bank, not this website. The real question is how reliable is reliable enough, and how you'd know. SRE answers it with a handful of acronyms that sound scary and are actually simple: measure the right thing (SLI), pick a target (SLO), and spend the gap between them wisely (error budget).

You will learn

  • The difference between an SLI, an SLO, an SLA and a KPI
  • What “three nines” really means in minutes
  • Error budgets and burn rate: how SLOs let teams move fast and stay reliable
  • Why percentiles beat averages
  • Measuring real SLIs from a web server log with awk

Four acronyms, one idea

Stands forIt's…Example
SLIService Level IndicatorA measurement of how the service behaves, as users see it. Usually “good events ÷ all events”.99.45% of requests yesterday returned a non-error status
SLOService Level ObjectiveA target for an SLI, over a time window. Set by the team.99.5% of requests succeed, measured over 30 days
SLAService Level AgreementA promise in a contract, with a penalty (usually money back) if it's broken. Set by lawyers and sales.“99.9% monthly uptime, or you get 10% credit”
KPIKey Performance IndicatorA business number leaders watch. Not reliability-specific.daily active users, sign-ups, revenue, tickets closed

An easy way to remember them: the Indicator is what you measure, the Objective is what you aim for, and the Agreement is what you've promised. SLOs are always stricter than SLAs, so the team gets a warning before a customer gets a refund. KPIs are the business's scoreboard. Reliability affects them (a slow checkout means fewer sales), which is why SREs care about them too.

Ops KPIs you'll hear

MTTD (mean time to detect), MTTR (mean time to restore), change failure rate (how many deploys cause a problem) and deployment frequency. The last two, with lead time and MTTR, are the four DORA metrics used to measure DevOps teams.

Choosing good SLIs

A good SLI measures what a user feels, not what a server feels. “CPU at 40%” isn't an SLI: no user has ever noticed your CPU. “Checkout pages load in under a second” is. Google's SRE book suggests starting from the four golden signals:

SignalQuestionExample SLI
LatencyHow long do requests take?% of requests served in ≤ 300 ms
TrafficHow much demand is there?requests per second
ErrorsHow many requests fail?% of requests with status ≥ 500
SaturationHow full is the system?disk 85% used, queue 150 deep

Traffic and saturation help you explain problems. Latency and errors are the ones users feel, so they make the best SLOs. Note that a 404 Not Found is usually the user's mistake (a bad link), so availability SLIs count only 5xx responses as failures.

How many nines?

Availability is usually written as “nines”. Each extra nine allows ten times less downtime, and it usually costs a lot more to reach:

TargetNameAllowed downtime per 30 days
99%two nines7 hours 12 minutes
99.5%two and a half3 hours 36 minutes
99.9%three nines43 minutes 12 seconds
99.99%four nines4 minutes 19 seconds
99.999%five nines26 seconds

The math is one line: 30 days × 24 × 60 × (1 − 0.999) = 43.2 minutes. Four nines leaves no time for a human to even wake up, so it needs automatic failover. Don't aim higher than your users need: if their home Wi-Fi is only 99% reliable, they'll never notice the difference between your 99.99% and 99.999%.

Error budgets

Here's the clever part. If the SLO is 99.5%, then 0.5% of requests are allowed to fail. That 0.5% is the error budget, and the team gets to spend it:

This ends the classic fight between developers (“ship faster!”) and operations (“change nothing!”). The data decides, and everyone agreed on the rules ahead of time, in an error budget policy.

Burn rate

Burn rate is how fast you're spending the budget. A burn rate of 1 uses exactly the whole budget in 30 days. A burn rate of 10 would use it all in 3 days. Good alerts fire on fast burn (page someone now) and slow burn (open a ticket), instead of on every single error.

Averages lie, percentiles don't

Suppose 99 requests take 50 ms and one takes 10 seconds. The average is about 150 ms, which describes nobody's actual experience. Percentiles do better: sort all the times, then pick a position.

In the practice log, the average is 109 ms and p95 is 195 ms, both comfortably fast. But p99 is 870 ms, and that's where the day's incident hides.

Measuring SLIs from a log

Before you have fancy monitoring (next lesson), your web server's access log already holds everything you need. In the practice log, each line ends with the response time in milliseconds:

198.51.100.8 - - [27/Sep/2026:00:00:26 +0000] "GET / HTTP/1.1" 200 2047 "-" "Mozilla/5.0" 88
$1 = client IP   $4 = time   $7 = path   $9 = status   $NF = the last field (ms)

This is the same combined log format as on your web servers. The response time is an extra you switch on: %D (microseconds) in Apache's LogFormat, or $request_time (seconds) in nginx's log_format.

Both families (awk is the same)
# availability SLI: % of requests that were NOT a 5xx
awk '$9 < 500 {good++} END {printf "%.2f%%\n", 100*good/NR}' access.log

# latency SLI: % of requests served in 300 ms or less
awk '$NF <= 300 {fast++} END {printf "%.2f%%\n", 100*fast/NR}' access.log

# percentiles: sort the times, then pick positions
awk '{print $NF}' access.log | sort -n |
  awk '{t[NR]=$1} END {print "p50", t[int(NR*0.5)], "p95", t[int(NR*0.95)], "p99", t[int(NR*0.99)]}'

# when did the errors happen? (characters 14-18 of $4 are HH:MM)
awk '$9 >= 500 {print substr($4, 14, 5)}' access.log | uniq -c

In awk, NR is the number of lines read so far, so in the END block it's the total. good++ counts, and printf formats the result to two decimal places.

gawk on Rocky, mawk on Ubuntu

Rocky's awk is GNU awk, and Ubuntu's is mawk (smaller and faster). Everything in this lesson works the same in both. They only differ in extras, like gawk's asort().

Practice: grade yesterday 📏

The ticket app's SLOs are in ~/slo.txt, and yesterday's log is in ~/logs/access.log. Did the team meet its targets? What happened, and how much budget did it cost?

Quick check

1. “99.9% of checkout requests succeed over 30 days” is…

2. Your SLA promises customers 99.9%. What should your internal SLO be?

3. The error budget has plenty left and the team wants to ship a risky new feature. What does an error budget policy say?

4. Average latency is 109 ms, but p99 is 870 ms. What does that tell you?

Finished the missions and the quiz? Mark it done to track your progress.