Linux SRE · Lesson 3 · 45 min

Monitoring & alerting with Prometheus

In the last two lessons you measured things by hand: vmstat during an incident, awk over yesterday's log. That works once. Monitoring does it all the time: it collects numbers every few seconds, keeps their history, draws graphs, and wakes someone up when a number crosses a line. You'll set up the most popular open-source monitoring system, Prometheus, from scratch, then write an alert and watch it fire.

You will learn

  • Metrics, logs and traces: the three kinds of monitoring data
  • How Prometheus works: exporters, scraping, labels and time series
  • Installing node_exporter and Prometheus as systemd services on Rocky and Ubuntu
  • PromQL: up, rate(), avg by, and turning numbers into percentages
  • Writing an alert rule, and what makes a good alert
  • The monitoring toolbox: Grafana, Alertmanager, Zabbix, Loki, OpenTelemetry and friends

Metrics, logs and traces

What it isGood forPopular tools
MetricsNumbers over time: CPU %, requests per second, free diskGraphs, trends, alerts. Cheap to keep for months.Prometheus, Grafana, Zabbix, Datadog, CloudWatch
LogsLines of text about individual eventsThe details: which request failed, and whyjournald, rsyslog, Loki, Elasticsearch/OpenSearch
TracesOne request's path through many services, with timings“Which of our 12 services made this request slow?”OpenTelemetry, Jaeger, Tempo

Metrics tell you that something is wrong, and logs and traces tell you why. Being able to answer new questions from this data is called observability. This lesson is about metrics.

How Prometheus works

  ┌──────────────┐   every 15 s: GET /metrics   ┌───────────────────┐
  │  Prometheus  │ ───────────────────────────▶ │ node_exporter     │  :9100
  │    :9090     │ ───────────────────────────▶ │ your app /metrics │  :8080
  │  stores the  │                              └───────────────────┘
  │  history,    │ ── alert rules ──▶ Alertmanager ──▶ email / Slack / PagerDuty
  │  runs PromQL │ ◀── queries ────── Grafana dashboards, you with curl

Metric types

TypeBehaves likeExampleHow you use it
Countera car's odometer: only goes up (and resets on restart)node_cpu_seconds_total, http_requests_totalNever graph it raw. Use rate() to get “per second”.
Gaugea speedometer: goes up and downnode_load1, node_filesystem_avail_bytesUse it as it is.
Histogramrequests sorted into buckets by size or timehttp_request_duration_seconds_buckethistogram_quantile(0.95, …) for percentiles

Install node_exporter and Prometheus

Both are single Go programs with no dependencies, so you install them the same way on every Linux: unpack the release tarball from prometheus.io/download, copy the program into /usr/local/bin, give it its own system user and run it with systemd. That's the same pattern you used for ourapp in Linux Sysadmin.

Rocky / RHEL

Prometheus isn't in Rocky's own repositories, so the release tarball (this lesson) or the official container image (quay.io/prometheus/prometheus) is the usual way. SELinux is happy with programs in /usr/local/bin.

Ubuntu / Debian

sudo apt install prometheus prometheus-node-exporter works too, and even writes the systemd units for you, but Ubuntu's versions lag behind upstream. The tarball gets you the current release on both families.

tar -xzf node_exporter-1.9.1.linux-amd64.tar.gz
tar -xzf prometheus-3.5.0.linux-amd64.tar.gz
sudo cp node_exporter-1.9.1.linux-amd64/node_exporter /usr/local/bin/
sudo cp prometheus-3.5.0.linux-amd64/prometheus prometheus-3.5.0.linux-amd64/promtool /usr/local/bin/

# one system user per service, with no login shell
sudo useradd -r -s /sbin/nologin node_exporter
sudo useradd -r -s /sbin/nologin prometheus

# the unit files (in ~/units in the practice terminal)
sudo cp units/*.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now node_exporter
curl -s localhost:9100/metrics | grep ^node_load

Prometheus needs a folder for its data (owned by prometheus, or it won't start) and a config file:

sudo mkdir -p /etc/prometheus/rules /var/lib/prometheus
sudo chown prometheus:prometheus /var/lib/prometheus
/etc/prometheus/prometheus.yml
global:
  scrape_interval: 15s        # how often to collect
  evaluation_interval: 15s    # how often to check alert rules

rule_files:
  - "rules/*.yml"

scrape_configs:
  - job_name: "prometheus"    # Prometheus watches itself too
    static_configs:
      - targets: ["localhost:9090"]

  - job_name: "node"
    static_configs:
      - targets: ["localhost:9100"]
promtool check config /etc/prometheus/prometheus.yml    # ALWAYS check before starting or reloading
sudo systemctl enable --now prometheus
curl -s localhost:9090/-/ready
YAML is picky about spaces

Indent with spaces, never tabs, and line up items at the same level exactly. A list item starts with - . Most “Prometheus won't start” problems are a misplaced space, and promtool check config tells you the line number.

Prometheus has a web UI on port 9090, with graphs, targets and alerts. Don't open 9090 (or 9100) to the whole internet: anyone could read your metrics. Use an SSH tunnel from your own computer instead, ssh -L 9090:localhost:9090 you@server, then browse to http://localhost:9090. In the terminal, the same data comes from the HTTP API, which you can read with curl and jq.

PromQL: asking questions

QuestionPromQL
Is every target up?up (1 = yes, 0 = down)
Load averagenode_load1
CPU busy %, per server100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])))
Disk free %node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"} * 100
Memory available %node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes * 100
Where is CPU time going?sum by (mode) (rate(node_cpu_seconds_total[5m]))
When will the disk be full?predict_linear(node_filesystem_avail_bytes{mountpoint="/"}[6h], 24*3600) < 0
curl -s 'localhost:9090/api/v1/query?query=up' | jq '.data.result[] | {job: .metric.job, up: .value[1]}'
promtool query instant http://localhost:9090 'node_load1'

Alerts that don't cry wolf

/etc/prometheus/rules/disk.yml
groups:
  - name: disk
    rules:
      - alert: DiskAlmostFull
        expr: node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"} * 100 < 20
        for: 5m                          # must stay true this long before it fires
        labels:
          severity: warning
        annotations:
          summary: "Disk almost full on {{ $labels.instance }}"
          description: "Only {{ $value | printf \"%.1f\" }}% free on /"
promtool check rules /etc/prometheus/rules/disk.yml
sudo systemctl reload prometheus          # re-reads config + rules, keeps running
curl -s localhost:9090/api/v1/alerts | jq

An alert is inactive, then pending (the condition is true, but the for: timer hasn't run out), then firing. for: stops a 10-second blip from paging anyone. Prometheus only decides which alerts fire. Alertmanager (a separate program) groups them, silences them during maintenance, and sends them to email, Slack, PagerDuty and so on.

Good alerting rules of thumb

Page a human only when users are hurting, or will be soon: an SLO burning fast, the site down, a disk that will be full before morning. Everything else becomes a ticket for working hours. Every page should be actionable, with a link to a runbook that says what to do. If people start ignoring an alert, fix it or delete it: alert fatigue is how real outages get missed.

The rest of the toolbox

ToolWhat it's for
GrafanaDashboards. It reads Prometheus (and much more) and draws graphs. Almost always paired with Prometheus.
AlertmanagerRouting, grouping and silencing alerts from Prometheus
Zabbix, Nagios / IcingaOlder all-in-one monitoring (checks, graphs, alerts, web UI). Common in companies with lots of servers, and fine choices.
NetdataVery detailed per-second dashboards for one machine, with almost no setup
Loki, Elasticsearch / OpenSearchCollect and search logs from all your servers in one place
OpenTelemetryThe standard way for apps to send metrics, logs and traces to any tool
Datadog, New Relic, CloudWatchPaid, hosted services: less to run yourself, and a monthly bill

Practice: watch the server, catch the disk 📈

The tarballs and unit files are in your home folder. Install both programs, point Prometheus at node_exporter, ask it some questions, then write an alert and make it fire.

Quick check

1. You graph node_cpu_seconds_total and it's a line that only goes up. What's wrong?

2. up{job="node"} is 0. What does it mean?

3. Why does the disk alert have for: 5m?

4. Which of these should page someone at 3 am?

Finished the missions and the quiz? Mark it done to track your progress.