Linux DevOps · Lesson 6 · 35 min

Deploy strategies: blue-green, rolling & rollback

Your CI pipeline is green and the new image is built. Now the scary part: putting it in front of real users. The worst way is to stop the old version, start the new one, and hope. This lesson is about the better ways: running old and new side by side, switching traffic in a second, and, most importantly, being able to switch back just as fast when something's wrong.

You will learn

  • Recreate, rolling, blue-green and canary deployments, and when to use each
  • Apache as a reverse proxy in front of containers: ProxyPass, configtest, graceful reload
  • Smoke tests and health checks before users see anything
  • Rollback vs roll forward
  • The hard parts: database changes and feature flags

Four ways to ship

StrategyHowDowntime?RollbackCost
RecreateStop old, start newYesSlow: redeploy the old oneCheapest
RollingReplace servers one (or a few) at a time behind a load balancerNoRoll back one at a timeNo extra servers
Blue-greenRun the new version (green) next to the old one (blue), test it, then switch all traffic at onceNoInstant: switch back to blueDouble, for a short time
CanarySend a small slice of traffic (1%, then 10%…) to the new version and watch the metricsNoFast: set the slice to 0%Needs good monitoring

They're named after real things. Canaries were carried into coal mines: if the bird got sick, the miners got out before they did. A canary release lets a few users hit a problem before everyone does. Rolling updates are what Kubernetes does by default. Blue-green is the easiest to understand and to practise, so that's this lesson's hands-on part.

Blue-green with Apache

                              ┌─────────────────────┐
  users ──▶ Apache :80 ──────▶│ blue   v1.0   :8081 │   ← live
             ProxyPass        └─────────────────────┘
             / → :8081        ┌─────────────────────┐
                              │ green  v2.0   :8082 │   ← being tested
                              └─────────────────────┘

Users only ever talk to Apache on port 80. Apache proxies each request to whichever container its ProxyPass line points at. Deploying means pointing it at the other color:

Rocky / RHEL
# /etc/httpd/conf.d/tickets.conf
ProxyPass        / http://127.0.0.1:8082/
ProxyPassReverse / http://127.0.0.1:8082/

sudo apachectl configtest
sudo systemctl reload httpd

SELinux normally stops Apache from connecting to other ports. Allow it once with sudo setsebool -P httpd_can_network_connect 1 (already done on the practice server).

Ubuntu / Debian
# /etc/apache2/sites-available/tickets.conf
ProxyPass        / http://127.0.0.1:8082/
ProxyPassReverse / http://127.0.0.1:8082/

sudo apache2ctl configtest
sudo systemctl reload apache2

The proxy modules must be enabled once: sudo a2enmod proxy proxy_http (already done on the practice server).

The same pattern works with nginx (proxy_pass + nginx -t + systemctl reload nginx), HAProxy, cloud load balancers, and Kubernetes Services. Only the file you edit changes.

Smoke tests and health checks

Before switching, test green directly on its own port, so users can't see it yet:

curl localhost:8082/health          # the app's own "I'm OK" endpoint
curl localhost:8082/                # the home page
curl localhost:8082/api/tickets     # and the things users actually rely on!

A smoke test is a quick check of the most important paths. Only testing /health is a classic mistake: the process is up, but the feature users need is broken. Good pipelines run smoke tests automatically and refuse to switch if they fail.

Rollback vs roll forward

When the new version misbehaves, you have two choices:

The rule from the incidents lesson applies: stop the pain first. Roll back, then fix calmly and deploy the fix through the normal pipeline. And don't delete blue the moment green is live. Keep it for a while so rollback stays instant.

The hard parts

Database changes

Blue and green usually share one database. If v2 renames a column, v1 breaks, and your instant rollback is gone. Teams use expand and contract: first add the new column while keeping the old one (both versions work), deploy, then remove the old column in a later release, once nobody needs to roll back.

Feature flags

Split deploying code from releasing a feature. Ship the new search hidden behind a flag, turn it on for staff, then 10% of users, then everyone. Turning it off again is a setting, not a deployment.

Practice: ship v2 without anyone noticing 🔵🟢

The ticket site runs as container blue (v1.0) behind Apache. Images for 2.0 and 2.0.1 are already built. Deploy 2.0 as green, find out it has a problem, roll back, then roll forward to 2.0.1.

Quick check

1. What makes rollback so fast with blue-green?

2. You edited ProxyPass to point at green, but curl localhost still shows v1.0. Why?

3. Green's /health says ok. Is it safe to switch?

4. Which strategy sends 5% of users to the new version first and watches the error rate?

Finished the missions and the quiz? Mark it done to track your progress.

Next up: Linux SRE · Keep it up, find out why, starting with “Performance triage: the 60-second checklist”.