Auto Scaling: servers that heal and grow
In the last lesson a broken server stayed broken until you fixed it. An Auto Scaling group (ASG) takes that job away from you. It keeps the number of servers you ask for, replaces any that die or fail health checks, and adds or removes servers as traffic changes. You describe how to build a server once, in a launch template, and AWS does the rest.
You will learn
- Launch templates: AMI, type, security group and base64 user data
- ASG sizes: min, max and desired capacity
- Self-healing with EC2 and ELB health checks, and the grace period
- Target tracking scaling policies, and the CloudWatch alarms behind them
- Reading
describe-scaling-activitiesto see what happened and why
Launch template: the recipe
A launch template holds everything run-instances needed in the EC2 lesson. Its user data must be base64-encoded: plain text squeezed into letters and digits so it survives inside JSON. (It's not encryption: base64 -d turns it straight back.)
base64 -w0 web.sh # -w0 = one long line, no wrapping
UD=$(base64 -w0 web.sh)
aws ec2 create-launch-template --launch-template-name web-lt --launch-template-data \
"{\"ImageId\":\"$AMI\",\"InstanceType\":\"t3.micro\",\"SecurityGroupIds\":[\"$SG\"],\"UserData\":\"$UD\"}"
The JSON sits inside double quotes so the shell fills in $AMI, $SG and $UD. That's why the inner quotes are written \". Templates have versions: change something with create-launch-template-version, and the ASG can use $Latest.
The group: how many, where, and who checks them
aws autoscaling create-auto-scaling-group --auto-scaling-group-name web-asg \ --launch-template LaunchTemplateName=web-lt,Version='$Latest' \ --min-size 2 --max-size 4 --desired-capacity 2 \ --vpc-zone-identifier "$SUBA,$SUBB" \ --target-group-arns $TG \ --health-check-type ELB --health-check-grace-period 60
| Setting | Meaning |
|---|---|
--min-size 2 | never fewer than 2, even at 3 a.m. |
--max-size 4 | never more than 4: your safety cap on the bill |
--desired-capacity 2 | how many right now. Scaling policies move this number between min and max |
--vpc-zone-identifier | subnets, comma-separated. The ASG spreads servers evenly across their AZs |
--target-group-arns | new servers are registered with the load balancer automatically |
--health-check-type ELB | replace a server if the load balancer says it's unhealthy, not only if EC2 says the VM is broken |
--health-check-grace-period 60 | give new servers 60 seconds to boot before judging them |
Scaling: let the numbers decide
The easiest policy is target tracking: "keep the average CPU around 50%". It works like a thermostat:
aws autoscaling put-scaling-policy --auto-scaling-group-name web-asg \
--policy-name cpu50 --policy-type TargetTrackingScaling \
--target-tracking-configuration \
'{"PredefinedMetricSpecification":{"PredefinedMetricType":"ASGAverageCPUUtilization"},"TargetValue":50.0}'
Behind the scenes, AWS creates two CloudWatch alarms: one that fires when CPU stays above 50% (scale out, quickly) and one for well below 50% (scale in, slowly, after about 15 minutes). You'll see both with aws cloudwatch describe-alarms. Don't edit them, since they belong to the policy. There are also step policies (your own alarm thresholds) and scheduled actions ("10 servers every weekday at 8:00").
Servers in an ASG are cattle: never fix one by hand, because the next replacement won't have your fix. Change the launch template (a new version), then start an instance refresh (aws autoscaling start-instance-refresh). The ASG replaces servers a few at a time while the load balancer keeps the site up. That's the same idea as the rolling deploys in the DevOps path.
Terminate its servers by hand and the ASG just launches new ones. That's its job! To really stop, set min/max/desired to 0, or delete the group with --force-delete, which terminates everything in it.
Practice: a self-healing, self-growing web tier 🌱
The load balancer web-alb, its listener and the target group web-tg from the last lesson are ready, with no servers behind them yet. web.sh is in your home folder. Let an ASG create the servers, kill one, then send a flood of visitors and watch it grow.
Quick check
1. An ASG has min 2, max 4, desired 2. You terminate one of its servers. What happens?
✓ It keeps "desired" true. To shrink, change desired (or let a scaling policy do it).
2. Why use --health-check-type ELB?
✓ A VM can be "running" while Apache is dead. The load balancer's health check notices that.
3. What stops a target tracking policy from launching 500 servers during a traffic spike?
✓ Max is your cap. Pick it with your wallet in mind.