Ansible in depth · Lesson 4 · 40 min

Fleets: inventories & rolling updates

Two servers you can update all at once. Twenty you shouldn't: if the new version is broken, all twenty break together. Real teams organise their machines into groups of groups, keep staging and production apart, and roll changes out a few servers at a time, checking each one and stopping at the first sign of trouble. Ansible does all of that with a handful of keywords.

You will learn

  • Groups of groups with [name:children], and ansible-inventory --graph
  • One inventory per environment: staging.ini and production.ini
  • Host patterns and --limit: web:!web3, web1,web2
  • Rolling updates: serial: and max_fail_percentage:
  • Health checks with uri + assert and delegate_to: localhost
  • Fixing one broken machine with an ad-hoc command, then finishing the rollout

Groups of groups

inventories/production.ini
[web]
web1 ansible_host=192.168.1.61     # Rocky
web2 ansible_host=192.168.1.62     # Ubuntu
web3 ansible_host=192.168.1.63     # Rocky

[db]
db1 ansible_host=192.168.1.71      # Ubuntu

[prod:children]                    # a group made of groups
web
db
$ ansible-inventory -i inventories/production.ini --graph
@all:
  |--@prod:
  |  |--@web:
  |  |  |--web1
  |  |  |--web2
  |  |  |--web3
  |  |--@db:
  |  |  |--db1

Now ansible prod -m ping reaches everything, ansible web … just the web servers, and group_vars/prod.yml could hold settings for the whole environment. Big companies go further, with inventory plugins that read the list of machines straight from their cloud account, so nobody maintains the file by hand.

Staging first, always

Keep one inventory per environment, and point the default at staging (inventory = inventories/staging.ini in ansible.cfg). That way a forgotten -i hits the safe environment, not production:

ansible-playbook rolling.yml                                  # staging (the default)
ansible-playbook -i inventories/production.ini rolling.yml    # production, on purpose
ansible-playbook -i inventories/production.ini rolling.yml -l web1        # just one machine
ansible-playbook -i inventories/production.ini rolling.yml -l 'web:!web3' # all web except web3

Rolling updates

- name: Rolling update of the web servers
  hosts: web
  become: true
  serial: 1                  # one server at a time (also: 2, or "25%")
  max_fail_percentage: 0     # any failure stops the rollout
  roles:
    - web
  post_tasks:
    - name: Health check from the control machine
      ansible.builtin.uri:
        url: "http://{{ inventory_hostname }}/health"
        return_content: true
      delegate_to: localhost  # run THIS task on the control machine, like a user would
      become: false
      register: health
    - name: Make sure the new version is live
      ansible.builtin.assert:
        that:
          - "'version ' ~ site_version in health.content"

With serial: 1, Ansible runs the whole play (roles, handlers, health check) on one server before it touches the next. If a server fails, max_fail_percentage: 0 stops everything: NO MORE HOSTS LEFT. The remaining servers keep running the old version, which is exactly what you want when the new one might be broken. Behind a load balancer, users never notice.

Health checks prove it, not just “changed”

A task saying changed means Ansible did something, not that the site works. uri + assert check what users actually see. Add the check to every rollout and you'll catch the problem on server one, not server twenty.

Two families, one rollout

The fleet mixes Rocky and Ubuntu. The role picks the right names from facts (apache_svc: "{{ 'httpd' if ansible_facts['os_family'] == 'RedHat' else 'apache2' }}" in its defaults), so the rollout doesn't care which family each server is.

When a rollout stops

  1. Don't panic: the servers that weren't touched are still fine.
  2. Look at the failed one with ad-hoc commands: ansible web2 -b -a 'apache2ctl configtest', -a 'journalctl -u apache2 -n 20'.
  3. Fix the cause (ideally in the role, so it can't come back), then rerun the same playbook. Idempotence means the finished servers just say ok, and the rollout continues where it stopped.

Practice: roll out version 2 🚚

web1, web2 and web3 all run version 1 of the site, and db1 runs the database. The role is ready to deploy version 2. Write the production inventory, test on staging, then roll out to production, and deal with what you find.

Quick check

1. With serial: 1 and max_fail_percentage: 0, web2 fails. What happens to web3?

2. Why does the health check use delegate_to: localhost?

3. Why make staging the default inventory in ansible.cfg?

Finished the missions and the quiz? Mark it done to track your progress.