Fleets: inventories & rolling updates
Two servers you can update all at once. Twenty you shouldn't: if the new version is broken, all twenty break together. Real teams organise their machines into groups of groups, keep staging and production apart, and roll changes out a few servers at a time, checking each one and stopping at the first sign of trouble. Ansible does all of that with a handful of keywords.
You will learn
- Groups of groups with
[name:children], andansible-inventory --graph - One inventory per environment:
staging.iniandproduction.ini - Host patterns and
--limit:web:!web3,web1,web2 - Rolling updates:
serial:andmax_fail_percentage: - Health checks with
uri+assertanddelegate_to: localhost - Fixing one broken machine with an ad-hoc command, then finishing the rollout
Groups of groups
[web] web1 ansible_host=192.168.1.61 # Rocky web2 ansible_host=192.168.1.62 # Ubuntu web3 ansible_host=192.168.1.63 # Rocky [db] db1 ansible_host=192.168.1.71 # Ubuntu [prod:children] # a group made of groups web db
$ ansible-inventory -i inventories/production.ini --graph @all: |--@prod: | |--@web: | | |--web1 | | |--web2 | | |--web3 | |--@db: | | |--db1
Now ansible prod -m ping reaches everything, ansible web … just the web servers, and group_vars/prod.yml could hold settings for the whole environment. Big companies go further, with inventory plugins that read the list of machines straight from their cloud account, so nobody maintains the file by hand.
Staging first, always
Keep one inventory per environment, and point the default at staging (inventory = inventories/staging.ini in ansible.cfg). That way a forgotten -i hits the safe environment, not production:
ansible-playbook rolling.yml # staging (the default) ansible-playbook -i inventories/production.ini rolling.yml # production, on purpose ansible-playbook -i inventories/production.ini rolling.yml -l web1 # just one machine ansible-playbook -i inventories/production.ini rolling.yml -l 'web:!web3' # all web except web3
Rolling updates
- name: Rolling update of the web servers hosts: web become: true serial: 1 # one server at a time (also: 2, or "25%") max_fail_percentage: 0 # any failure stops the rollout roles: - web post_tasks: - name: Health check from the control machine ansible.builtin.uri: url: "http://{{ inventory_hostname }}/health" return_content: true delegate_to: localhost # run THIS task on the control machine, like a user would become: false register: health - name: Make sure the new version is live ansible.builtin.assert: that: - "'version ' ~ site_version in health.content"
With serial: 1, Ansible runs the whole play (roles, handlers, health check) on one server before it touches the next. If a server fails, max_fail_percentage: 0 stops everything: NO MORE HOSTS LEFT. The remaining servers keep running the old version, which is exactly what you want when the new one might be broken. Behind a load balancer, users never notice.
A task saying changed means Ansible did something, not that the site works. uri + assert check what users actually see. Add the check to every rollout and you'll catch the problem on server one, not server twenty.
The fleet mixes Rocky and Ubuntu. The role picks the right names from facts (apache_svc: "{{ 'httpd' if ansible_facts['os_family'] == 'RedHat' else 'apache2' }}" in its defaults), so the rollout doesn't care which family each server is.
When a rollout stops
- Don't panic: the servers that weren't touched are still fine.
- Look at the failed one with ad-hoc commands:
ansible web2 -b -a 'apache2ctl configtest',-a 'journalctl -u apache2 -n 20'. - Fix the cause (ideally in the role, so it can't come back), then rerun the same playbook. Idempotence means the finished servers just say ok, and the rollout continues where it stopped.
Practice: roll out version 2 🚚
web1, web2 and web3 all run version 1 of the site, and db1 runs the database. The role is ready to deploy version 2. Write the production inventory, test on staging, then roll out to production, and deal with what you find.
Quick check
1. With serial: 1 and max_fail_percentage: 0, web2 fails. What happens to web3?
✓ One broken server instead of three.
2. Why does the health check use delegate_to: localhost?
✓ Test what users experience.
3. Why make staging the default inventory in ansible.cfg?
✓ Make the dangerous thing the one you have to type on purpose.