CloudWatch: metrics, alarms & billing alerts
You can't stare at servers all day, and you shouldn't have to. CloudWatch collects numbers (metrics) from everything you run, keeps logs, and fires alarms when a number crosses a line. Alarms send messages through SNS, such as an email to you. The same tools protect your wallet: a billing alarm and a budget are the first thing to set up in any real account.
You will learn
- Metrics: namespace, name, dimensions, period, statistic
list-metricsandget-metric-statistics(withdate -dfor time ranges)- SNS topics and email subscriptions, including confirming them
- Alarms: thresholds, evaluation periods, the three states, and actions
- A billing alarm, an AWS Budget, and CloudWatch Logs retention
What a metric is
A metric is a number over time, with a name tag. The SRE path's Prometheus did the same job. Here's how CloudWatch names things:
| Part | Example |
|---|---|
| Namespace: which service | AWS/EC2, AWS/ApplicationELB, AWS/Billing, or your own like CHT/App |
| Metric name | CPUUtilization, RequestCount, EstimatedCharges |
| Dimensions: which one | InstanceId=i-0abc…, AutoScalingGroupName=web-asg |
| Period + statistic: how to summarise | Average / Maximum / Sum over 60 or 300 seconds |
EC2 sends CPU, network and disk-operation metrics every 5 minutes for free (every minute if you pay for detailed monitoring). Memory and disk space are not included, because AWS can't see inside your server. For those you install the CloudWatch agent on the instance.
aws cloudwatch list-metrics --namespace AWS/EC2 --dimensions Name=InstanceId,Value=$ID aws cloudwatch get-metric-statistics --namespace AWS/EC2 --metric-name CPUUtilization \ --dimensions Name=InstanceId,Value=$ID --statistics Average Maximum --period 300 \ --start-time $(date -u -d '-1 hour' +%Y-%m-%dT%H:%M:%SZ) --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ)
date -u -d '-1 hour' works out "one hour ago" in UTC, and the +%Y-%m-%dT… format is the ISO time AWS expects.
SNS: how alarms reach you
TOPIC=$(aws sns create-topic --name cht-alerts --query TopicArn --output text) aws sns subscribe --topic-arn $TOPIC --protocol email --notification-endpoint you@example.com
A topic is a mailing list for machines: publish once, and every subscription gets a copy (email, SMS, a web hook, a queue, a Lambda function…). Email subscriptions stay PendingConfirmation until someone clicks the link AWS sends, so nobody can sign you up for spam. In the playground, email arrives in ~/inbox, and you "click" with curl.
Alarms
aws cloudwatch put-metric-alarm --alarm-name cpu-high \ --namespace AWS/EC2 --metric-name CPUUtilization --dimensions Name=InstanceId,Value=$ID \ --statistic Average --period 60 --evaluation-periods 2 \ --threshold 70 --comparison-operator GreaterThanThreshold \ --alarm-actions $TOPIC
"If the average CPU over 60 seconds is above 70, twice in a row, go to ALARM and tell the topic." An alarm is always in one of three states:
- OK: the metric is on the right side of the line.
- ALARM: it crossed the line for enough periods. Actions run on the change into this state, not over and over.
- INSUFFICIENT_DATA: not enough numbers yet (just created, or the thing stopped reporting).
--treat-missing-datadecides what silence means.
Actions can do more than send a message: arn:aws:automate:us-east-1:ec2:stop stops the instance, and Auto Scaling policies are alarm actions too. That's how the last lesson's scaling worked. aws cloudwatch set-alarm-state forces a state for testing, so you can check the email arrives without breaking anything.
From the SRE path: "CPU is 80%" isn't a problem by itself, but "5% of requests fail" is. CloudWatch has the load balancer's HTTPCode_Target_5XX_Count and TargetResponseTime. Alarms on those are closer to your SLOs than CPU alarms.
Protect your wallet
AWS won't stop you from spending. It will tell you, if you ask. Do both of these on day one of a real account:
- A billing alarm on
AWS/Billing EstimatedCharges(the month-to-date bill, inCurrency=USD). Billing metrics exist only in us-east-1, whatever region you work in, and you must first switch on "Receive Billing Alerts" in the billing console's preferences. - AWS Budgets: "tell me when this month's cost goes over $10, or is forecast to". The first budgets are free.
aws cloudwatch put-metric-alarm --region us-east-1 --alarm-name billing-over-10 \ --namespace AWS/Billing --metric-name EstimatedCharges --dimensions Name=Currency,Value=USD \ --statistic Maximum --period 21600 --evaluation-periods 1 \ --threshold 10 --comparison-operator GreaterThanThreshold --alarm-actions $TOPIC aws budgets create-budget --account-id 123456789012 --budget file://budget.json \ --notifications-with-subscribers file://notify.json
Logs
CloudWatch Logs stores log lines in log groups (one per app, e.g. /cht/web), each with streams (one per server). The CloudWatch agent or your app sends them, and aws logs tail /cht/web --follow watches them live. New log groups keep everything forever by default, and you pay for every GB stored. Always set a retention:
aws logs put-retention-policy --log-group-name /cht/web --retention-in-days 14
Practice: get woken up (by email) 🔔
One web server, web-1, is running, and the key to log in to it is ~/cht-key.pem. Watch its CPU, set up an email alarm, then make the CPU busy on purpose and see the alarm fire. Finish with the money alarms.
Quick check
1. You made a billing alarm in eu-west-2 and it stays INSUFFICIENT_DATA forever. Why?
✓ And switch on "Receive Billing Alerts" in the billing preferences first.
2. Your alarm went to ALARM, but no email came. What's the most likely reason?
✓ aws sns list-subscriptions-by-topic shows PendingConfirmation until the link is clicked.
3. Which EC2 number does CloudWatch not have unless you install the CloudWatch agent?
✓ AWS sees the VM from outside. Memory and disk space need an agent inside.