Monitoring & alerting with Prometheus
In the last two lessons you measured things by hand: vmstat during an incident, awk over yesterday's log. That works once. Monitoring does it all the time: it collects numbers every few seconds, keeps their history, draws graphs, and wakes someone up when a number crosses a line. You'll set up the most popular open-source monitoring system, Prometheus, from scratch, then write an alert and watch it fire.
You will learn
- Metrics, logs and traces: the three kinds of monitoring data
- How Prometheus works: exporters, scraping, labels and time series
- Installing node_exporter and Prometheus as systemd services on Rocky and Ubuntu
- PromQL:
up,rate(),avg by, and turning numbers into percentages - Writing an alert rule, and what makes a good alert
- The monitoring toolbox: Grafana, Alertmanager, Zabbix, Loki, OpenTelemetry and friends
Metrics, logs and traces
| What it is | Good for | Popular tools | |
|---|---|---|---|
| Metrics | Numbers over time: CPU %, requests per second, free disk | Graphs, trends, alerts. Cheap to keep for months. | Prometheus, Grafana, Zabbix, Datadog, CloudWatch |
| Logs | Lines of text about individual events | The details: which request failed, and why | journald, rsyslog, Loki, Elasticsearch/OpenSearch |
| Traces | One request's path through many services, with timings | “Which of our 12 services made this request slow?” | OpenTelemetry, Jaeger, Tempo |
Metrics tell you that something is wrong, and logs and traces tell you why. Being able to answer new questions from this data is called observability. This lesson is about metrics.
How Prometheus works
┌──────────────┐ every 15 s: GET /metrics ┌───────────────────┐ │ Prometheus │ ───────────────────────────▶ │ node_exporter │ :9100 │ :9090 │ ───────────────────────────▶ │ your app /metrics │ :8080 │ stores the │ └───────────────────┘ │ history, │ ── alert rules ──▶ Alertmanager ──▶ email / Slack / PagerDuty │ runs PromQL │ ◀── queries ────── Grafana dashboards, you with curl
- An exporter is a small program that turns something into metrics on an HTTP page. node_exporter does it for a Linux machine. There are exporters for MySQL, nginx, Apache and hundreds more, and many apps expose
/metricsthemselves. - Prometheus pulls (“scrapes”) every target on a schedule. If a target doesn't answer, Prometheus records
up = 0, which is itself the most useful alert there is. - Every series has a name and labels:
node_cpu_seconds_total{cpu="0", mode="idle"}. Labels are how you filter and group.
Metric types
| Type | Behaves like | Example | How you use it |
|---|---|---|---|
| Counter | a car's odometer: only goes up (and resets on restart) | node_cpu_seconds_total, http_requests_total | Never graph it raw. Use rate() to get “per second”. |
| Gauge | a speedometer: goes up and down | node_load1, node_filesystem_avail_bytes | Use it as it is. |
| Histogram | requests sorted into buckets by size or time | http_request_duration_seconds_bucket | histogram_quantile(0.95, …) for percentiles |
Install node_exporter and Prometheus
Both are single Go programs with no dependencies, so you install them the same way on every Linux: unpack the release tarball from prometheus.io/download, copy the program into /usr/local/bin, give it its own system user and run it with systemd. That's the same pattern you used for ourapp in Linux Sysadmin.
Prometheus isn't in Rocky's own repositories, so the release tarball (this lesson) or the official container image (quay.io/prometheus/prometheus) is the usual way. SELinux is happy with programs in /usr/local/bin.
sudo apt install prometheus prometheus-node-exporter works too, and even writes the systemd units for you, but Ubuntu's versions lag behind upstream. The tarball gets you the current release on both families.
tar -xzf node_exporter-1.9.1.linux-amd64.tar.gz tar -xzf prometheus-3.5.0.linux-amd64.tar.gz sudo cp node_exporter-1.9.1.linux-amd64/node_exporter /usr/local/bin/ sudo cp prometheus-3.5.0.linux-amd64/prometheus prometheus-3.5.0.linux-amd64/promtool /usr/local/bin/ # one system user per service, with no login shell sudo useradd -r -s /sbin/nologin node_exporter sudo useradd -r -s /sbin/nologin prometheus # the unit files (in ~/units in the practice terminal) sudo cp units/*.service /etc/systemd/system/ sudo systemctl daemon-reload sudo systemctl enable --now node_exporter curl -s localhost:9100/metrics | grep ^node_load
Prometheus needs a folder for its data (owned by prometheus, or it won't start) and a config file:
sudo mkdir -p /etc/prometheus/rules /var/lib/prometheus sudo chown prometheus:prometheus /var/lib/prometheus
global: scrape_interval: 15s # how often to collect evaluation_interval: 15s # how often to check alert rules rule_files: - "rules/*.yml" scrape_configs: - job_name: "prometheus" # Prometheus watches itself too static_configs: - targets: ["localhost:9090"] - job_name: "node" static_configs: - targets: ["localhost:9100"]
promtool check config /etc/prometheus/prometheus.yml # ALWAYS check before starting or reloading
sudo systemctl enable --now prometheus
curl -s localhost:9090/-/ready
Indent with spaces, never tabs, and line up items at the same level exactly. A list item starts with - . Most “Prometheus won't start” problems are a misplaced space, and promtool check config tells you the line number.
Prometheus has a web UI on port 9090, with graphs, targets and alerts. Don't open 9090 (or 9100) to the whole internet: anyone could read your metrics. Use an SSH tunnel from your own computer instead, ssh -L 9090:localhost:9090 you@server, then browse to http://localhost:9090. In the terminal, the same data comes from the HTTP API, which you can read with curl and jq.
PromQL: asking questions
| Question | PromQL |
|---|---|
| Is every target up? | up (1 = yes, 0 = down) |
| Load average | node_load1 |
| CPU busy %, per server | 100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))) |
| Disk free % | node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"} * 100 |
| Memory available % | node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes * 100 |
| Where is CPU time going? | sum by (mode) (rate(node_cpu_seconds_total[5m])) |
| When will the disk be full? | predict_linear(node_filesystem_avail_bytes{mountpoint="/"}[6h], 24*3600) < 0 |
{mode="idle"}filters by label.=~matches a regex, and!=excludes.[5m]means “the last 5 minutes of samples”, andrate()turns that counter history into a per-second speed. Idle seconds per second is the fraction of time the CPU was idle.avg by (instance)merges the 4 CPUs into one number per server.sum,max,minandcountwork the same way.
curl -s 'localhost:9090/api/v1/query?query=up' | jq '.data.result[] | {job: .metric.job, up: .value[1]}'
promtool query instant http://localhost:9090 'node_load1'
Alerts that don't cry wolf
groups:
- name: disk
rules:
- alert: DiskAlmostFull
expr: node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"} * 100 < 20
for: 5m # must stay true this long before it fires
labels:
severity: warning
annotations:
summary: "Disk almost full on {{ $labels.instance }}"
description: "Only {{ $value | printf \"%.1f\" }}% free on /"promtool check rules /etc/prometheus/rules/disk.yml
sudo systemctl reload prometheus # re-reads config + rules, keeps running
curl -s localhost:9090/api/v1/alerts | jq
An alert is inactive, then pending (the condition is true, but the for: timer hasn't run out), then firing. for: stops a 10-second blip from paging anyone. Prometheus only decides which alerts fire. Alertmanager (a separate program) groups them, silences them during maintenance, and sends them to email, Slack, PagerDuty and so on.
Page a human only when users are hurting, or will be soon: an SLO burning fast, the site down, a disk that will be full before morning. Everything else becomes a ticket for working hours. Every page should be actionable, with a link to a runbook that says what to do. If people start ignoring an alert, fix it or delete it: alert fatigue is how real outages get missed.
The rest of the toolbox
| Tool | What it's for |
|---|---|
| Grafana | Dashboards. It reads Prometheus (and much more) and draws graphs. Almost always paired with Prometheus. |
| Alertmanager | Routing, grouping and silencing alerts from Prometheus |
| Zabbix, Nagios / Icinga | Older all-in-one monitoring (checks, graphs, alerts, web UI). Common in companies with lots of servers, and fine choices. |
| Netdata | Very detailed per-second dashboards for one machine, with almost no setup |
| Loki, Elasticsearch / OpenSearch | Collect and search logs from all your servers in one place |
| OpenTelemetry | The standard way for apps to send metrics, logs and traces to any tool |
| Datadog, New Relic, CloudWatch | Paid, hosted services: less to run yourself, and a monthly bill |
Practice: watch the server, catch the disk 📈
The tarballs and unit files are in your home folder. Install both programs, point Prometheus at node_exporter, ask it some questions, then write an alert and make it fire.
Quick check
1. You graph node_cpu_seconds_total and it's a line that only goes up. What's wrong?
✓ Counters only go up. rate() shows how fast they're climbing, and that's the useful part.
2. up{job="node"} is 0. What does it mean?
✓ /api/v1/targets shows the exact error (lastError), for example “connection refused”.
3. Why does the disk alert have for: 5m?
✓ Until then the alert is “pending”.
4. Which of these should page someone at 3 am?
✓ Page on user pain (symptoms). The other two are tickets for the morning.