AWS basics · Lesson 9 · 40 min

CloudWatch: metrics, alarms & billing alerts

You can't stare at servers all day, and you shouldn't have to. CloudWatch collects numbers (metrics) from everything you run, keeps logs, and fires alarms when a number crosses a line. Alarms send messages through SNS, such as an email to you. The same tools protect your wallet: a billing alarm and a budget are the first thing to set up in any real account.

You will learn

  • Metrics: namespace, name, dimensions, period, statistic
  • list-metrics and get-metric-statistics (with date -d for time ranges)
  • SNS topics and email subscriptions, including confirming them
  • Alarms: thresholds, evaluation periods, the three states, and actions
  • A billing alarm, an AWS Budget, and CloudWatch Logs retention

What a metric is

A metric is a number over time, with a name tag. The SRE path's Prometheus did the same job. Here's how CloudWatch names things:

PartExample
Namespace: which serviceAWS/EC2, AWS/ApplicationELB, AWS/Billing, or your own like CHT/App
Metric nameCPUUtilization, RequestCount, EstimatedCharges
Dimensions: which oneInstanceId=i-0abc…, AutoScalingGroupName=web-asg
Period + statistic: how to summariseAverage / Maximum / Sum over 60 or 300 seconds

EC2 sends CPU, network and disk-operation metrics every 5 minutes for free (every minute if you pay for detailed monitoring). Memory and disk space are not included, because AWS can't see inside your server. For those you install the CloudWatch agent on the instance.

aws cloudwatch list-metrics --namespace AWS/EC2 --dimensions Name=InstanceId,Value=$ID
aws cloudwatch get-metric-statistics --namespace AWS/EC2 --metric-name CPUUtilization \
  --dimensions Name=InstanceId,Value=$ID --statistics Average Maximum --period 300 \
  --start-time $(date -u -d '-1 hour' +%Y-%m-%dT%H:%M:%SZ) --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ)

date -u -d '-1 hour' works out "one hour ago" in UTC, and the +%Y-%m-%dT… format is the ISO time AWS expects.

SNS: how alarms reach you

TOPIC=$(aws sns create-topic --name cht-alerts --query TopicArn --output text)
aws sns subscribe --topic-arn $TOPIC --protocol email --notification-endpoint you@example.com

A topic is a mailing list for machines: publish once, and every subscription gets a copy (email, SMS, a web hook, a queue, a Lambda function…). Email subscriptions stay PendingConfirmation until someone clicks the link AWS sends, so nobody can sign you up for spam. In the playground, email arrives in ~/inbox, and you "click" with curl.

Alarms

aws cloudwatch put-metric-alarm --alarm-name cpu-high \
  --namespace AWS/EC2 --metric-name CPUUtilization --dimensions Name=InstanceId,Value=$ID \
  --statistic Average --period 60 --evaluation-periods 2 \
  --threshold 70 --comparison-operator GreaterThanThreshold \
  --alarm-actions $TOPIC

"If the average CPU over 60 seconds is above 70, twice in a row, go to ALARM and tell the topic." An alarm is always in one of three states:

Actions can do more than send a message: arn:aws:automate:us-east-1:ec2:stop stops the instance, and Auto Scaling policies are alarm actions too. That's how the last lesson's scaling worked. aws cloudwatch set-alarm-state forces a state for testing, so you can check the email arrives without breaking anything.

Alert on what users feel

From the SRE path: "CPU is 80%" isn't a problem by itself, but "5% of requests fail" is. CloudWatch has the load balancer's HTTPCode_Target_5XX_Count and TargetResponseTime. Alarms on those are closer to your SLOs than CPU alarms.

Protect your wallet

AWS won't stop you from spending. It will tell you, if you ask. Do both of these on day one of a real account:

aws cloudwatch put-metric-alarm --region us-east-1 --alarm-name billing-over-10 \
  --namespace AWS/Billing --metric-name EstimatedCharges --dimensions Name=Currency,Value=USD \
  --statistic Maximum --period 21600 --evaluation-periods 1 \
  --threshold 10 --comparison-operator GreaterThanThreshold --alarm-actions $TOPIC

aws budgets create-budget --account-id 123456789012 --budget file://budget.json \
  --notifications-with-subscribers file://notify.json

Logs

CloudWatch Logs stores log lines in log groups (one per app, e.g. /cht/web), each with streams (one per server). The CloudWatch agent or your app sends them, and aws logs tail /cht/web --follow watches them live. New log groups keep everything forever by default, and you pay for every GB stored. Always set a retention:

aws logs put-retention-policy --log-group-name /cht/web --retention-in-days 14

Practice: get woken up (by email) 🔔

One web server, web-1, is running, and the key to log in to it is ~/cht-key.pem. Watch its CPU, set up an email alarm, then make the CPU busy on purpose and see the alarm fire. Finish with the money alarms.

Quick check

1. You made a billing alarm in eu-west-2 and it stays INSUFFICIENT_DATA forever. Why?

2. Your alarm went to ALARM, but no email came. What's the most likely reason?

3. Which EC2 number does CloudWatch not have unless you install the CloudWatch agent?

Finished the missions and the quiz? Mark it done to track your progress.