Load balancers: ELB and ALB
One server is one point of failure. Two servers need something in front that spreads visitors between them and stops sending anyone to a server that's sick. That's a load balancer. On AWS it's Elastic Load Balancing, and for websites you'll almost always use an Application Load Balancer (ALB). You'll build one in front of two servers in two availability zones, and watch it route around a broken one.
You will learn
- The ALB family: ALB, NLB and the old Classic, and when to use each
- The three pieces: load balancer, listener, target group
- Health checks and target states:
initial,healthy,unhealthy,unused - Chaining security groups: internet → ALB → servers
- What 502, 503 and 504 from a load balancer mean
Which load balancer?
| Type | Works at | Good for |
|---|---|---|
| Application (ALB) | HTTP/HTTPS (layer 7): sees URLs, host names, headers | websites and APIs. Can route /api to one group and / to another, and do HTTPS for you |
| Network (NLB) | TCP/UDP (layer 4): just passes connections along | non-HTTP protocols, huge traffic, fixed IP addresses |
| Classic | both, the old way | nothing new. It's being retired |
The three pieces
internet
│ http://web-alb-123.us-east-1.elb.amazonaws.com
▼
┌─── Load balancer (alb-sg: port 80 from anywhere) ────────┐
│ in subnet public-a (us-east-1a) and public-b (1b) │
│ │
│ Listener HTTP :80 → forward to target group │
└────────────────────────────┬─────────────────────────────┘
▼
Target group web-tg health check: GET / every 30 s
├── web-1 (us-east-1a) healthy ✓
└── web-2 (us-east-1b) healthy ✓ (web-sg: port 80 only from alb-sg)
- The load balancer needs subnets in at least two AZs. AWS runs it on its own nodes in each one, so it keeps working if one AZ goes down. You reach it by its DNS name, never by IP: its IPs change.
- A listener waits on a port and protocol (HTTP 80, HTTPS 443) and has rules for what to do: forward to a target group, redirect HTTP to HTTPS, or return a fixed response.
- A target group is a list of targets (instances, IPs or Lambda functions) plus a health check. Only healthy targets get traffic.
Building it
TG=$(aws elbv2 create-target-group --name web-tg --protocol HTTP --port 80 \ --vpc-id $VPC --health-check-path / --query 'TargetGroups[0].TargetGroupArn' --output text) aws elbv2 register-targets --target-group-arn $TG --targets Id=$WEB1 Id=$WEB2 LB=$(aws elbv2 create-load-balancer --name web-alb --subnets $SUBA $SUBB \ --security-groups $ALBSG --query 'LoadBalancers[0].LoadBalancerArn' --output text) aws elbv2 create-listener --load-balancer-arn $LB --protocol HTTP --port 80 \ --default-actions Type=forward,TargetGroupArn=$TG aws elbv2 wait load-balancer-available --load-balancer-arns $LB aws elbv2 describe-target-health --target-group-arn $TG
Load balancers, target groups and listeners are named by ARNs, not short IDs, so keeping them in variables really helps.
Security groups that point at each other
The best setup: the internet can reach the load balancer, and only the load balancer can reach the servers. Instead of IP addresses, the servers' group allows traffic from the load balancer's security group:
aws ec2 authorize-security-group-ingress --group-id $WEBSG \ --protocol tcp --port 80 --source-group $ALBSG
That rule keeps working however many load balancer nodes there are and whatever their IPs are. Nobody on the internet can reach the servers directly, so you can't be attacked around the load balancer.
Reading target health
| State / reason | Means | Look at |
|---|---|---|
unused Target.NotInUse | the target group isn't attached to a listener yet | create the listener |
initial | just registered, first checks running | wait a little |
unhealthy Target.Timeout | the health check never got an answer: the packets were dropped | the servers' security group |
unhealthy Target.FailedHealthChecks | connection refused, or nothing running | is the web server running on the instance? |
unhealthy Target.ResponseCodeMismatch | it answered, but not with 200 | the health check path, and the app |
unused Target.InvalidState | the instance is stopped | start it (or let Auto Scaling replace it) |
And the errors a visitor might see from the load balancer itself (look for Server: awselb/2.0):
- 503 Service Unavailable: no healthy targets at all.
- 504 Gateway Timeout: a target didn't answer in time. Very often a security group again.
- 502 Bad Gateway: a target answered with something broken, or closed the connection.
With a free certificate from AWS Certificate Manager, an HTTPS listener on port 443 handles all the encryption. The servers behind it can keep speaking plain HTTP inside the VPC. Add an HTTP listener that redirects to HTTPS and you're done.
An ALB costs about 2.25 cents an hour (roughly $16 a month) plus a small charge for traffic, even with no visitors. Delete practice load balancers when you're finished.
Practice: two servers, one front door ⚖️
Two web servers are already running: web-1 in us-east-1a and web-2 in us-east-1b. Their security group web-sg allows SSH from you, but not port 80 from anywhere, on purpose. Put an ALB in front, find out why it's unhappy, fix it the right way, then break a server and watch the ALB cope.
Quick check
1. Both targets show unhealthy with reason Target.Timeout. What's the most likely cause?
✓ Timeout = dropped packets = security group. Allow port 80 from the ALB's security group.
2. Why must an ALB have subnets in at least two availability zones?
✓ High availability is the whole point. Put your servers in two AZs as well.
3. Visitors get 503 from awselb/2.0. What does that tell you?
✓ Check describe-target-health. It tells you why each target is out.