Lesson 024 · Ansible and RHCE Automation Learning Path

Ansible RHCE Practical Capstone and Troubleshooting

· Published · 6 min read

Labelled Ansible control node workflow through inventory playbook modules SSH managed nodes idempotent second run verification failure evidence and safe rollback

The capstone starts from fresh RHEL nodes and requires inventory, roles, Vault, repositories, services, storage, security policy, validation, recovery and retained evidence under time pressure. The useful question is not whether the playbook finished; it is whether another operator can explain the selected hosts, inputs, module decisions, changes, failures and rollback from retained evidence.

Inherited lab checkpoint and starting evidence

Begin with the accepted checkpoint from the preceding lesson. Confirm both managed nodes answer the inventory and preserve the current project commit before changing this layer.

pwd
ansible --version
ansible-config dump --only-changed
ansible-inventory --graph
git status --short

Capture this output before changing the objective-mapped capstone. Keep host keys, vault passwords, private keys and tokens out of terminal transcripts and Git. The examples use control.example.test, nodea.example.test and nodeb.example.test; replace them only with identities verified in your own inventory.

Understand the objective-mapped capstone

QuestionOperator decisionEvidence
ScopeWhich hosts and groups should receive the change?Inventory graph and explicit limit
InputWhere does each value originate?Variable inspection without secret disclosure
StateWhich module expresses the required result?Module documentation and diff
FailureWhat must stop, continue or recover?Recap, registered result and managed-node logs
PersistenceDoes the result survive service restart or reboot?Second run and client-side acceptance

Prepare the project safely

cd ~/ansible-lab
git status --short
ansible-inventory -i inventories/lab.ini --graph
ansible all -i inventories/lab.ini -m ansible.builtin.ping --limit nodea.example.test

Work in a dedicated Git repository, inspect configuration precedence and commit no generated secrets. Use a named inventory and an explicit limit until host selection is proven.

Build the complete working example

# capstone acceptance
ansible-inventory -i inventories/exam.yml --graph
ansible-playbook --syntax-check -i inventories/exam.yml playbooks/site.yml
ansible-playbook --check --diff --limit nodea.example.test -i inventories/exam.yml playbooks/site.yml
ansible-playbook -i inventories/exam.yml playbooks/site.yml
ansible-playbook -i inventories/exam.yml playbooks/site.yml
curl --fail http://nodea.example.test:8080/health
ssh nodea.example.test 'sudo restorecon -RF /srv/app; systemctl --failed'

Build small roles with explicit ownership. Validate each objective as it is completed, save time for a clean second run and reboot test, and diagnose failures from the lowest failed layer.

Run, inspect and repeat

ansible-playbook --syntax-check -i inventories/lab.ini playbooks/site.yml
ansible-playbook --check --diff -i inventories/lab.ini playbooks/site.yml
ansible-playbook -i inventories/lab.ini playbooks/site.yml
ansible-playbook -i inventories/lab.ini playbooks/site.yml

The first run may report a controlled change. The second run should normally report changed=0 for the same desired state. If it changes again, identify the non-idempotent task rather than accepting noisy automation as normal.

Interpret the execution result

SignalHealthy meaningWhat a different result means
okTask inspected state and required no changeConfirm this was the intended host and state
changedModule made a declared changeReview diff and handler notification
failedTask could not establish its contractRead module message and managed-node evidence
unreachableConnection or transport failed before task executionCheck inventory, SSH, host key, route and Python
rescued or ignoredPlay continued under explicit failure policyEnsure the exception is visible and owned

Read the recap as a starting point, not as the acceptance test. A green play can still select the wrong host, install an unintended version, expose a service on the wrong interface or leave a change that disappears after reboot. Tie each requirement to evidence from the managed node and, where practical, to a client-side test. Keep the command, relevant output, inventory limit and Git revision together so another administrator can reproduce the decision.

When a run fails, resist changing several layers at once. First confirm inventory selection and transport, then privilege, input data, module arguments, managed-node state and finally the application response. Make one attributable correction and rerun the smallest safe scope. This preserves the causal evidence that disappears when shell commands, manual edits and repeated full-fleet runs are mixed together.

Practical use cases

Use caseImplementation choiceAcceptance
Complete estateApply roles to web and database groups with serial acceptanceEvery stated objective passes after reboot and second run
Narrow rolloutUse --limit and serial execution before the complete groupOnly intended hosts change and availability remains inside its budget
Dependency outageStop the named lab dependency and retain the failed resultFailure is visible, bounded and recoverable without manual drift

A timed run ends with changed tasks on every pass

The play reaches the apparent end state, but shell commands, timestamps and unconditional restarts remain noisy. Replace each with state modules or exact change tests, then rerun until only deliberate changes remain.

Troubleshoot by symptom

SymptomInspect firstCorrection
Several symptoms appear togetherFind the earliest failed task and inspect its layerCorrect root cause, rerun from a clean checkpoint
Host is unreachableInventory variables, DNS, SSH host key and PythonRepair connection ownership before changing the play
Second run changes againDiff, volatile input and task semanticsReplace imperative work with a stable desired-state test

Unsafe shortcuts and recovery boundaries

  • Unsafe: copying an unknown solution sacrifices understanding and may target the wrong environment.
  • Unsafe: making broad manual repairs creates drift the final play cannot reproduce.
  • Unsafe: skipping reboot validation leaves mounts, services and policy persistence unproven.

Production operation and rollback

Store the reviewed project in Git, pin external content, separate inventory data from secrets and promote the same commit through environments. Monitor unreachable, failed, rescued, ignored and changed results separately.

ansible-playbook --syntax-check -i inventories/lab.ini playbooks/rollback.yml
ansible-playbook --check --diff --limit nodea.example.test -i inventories/lab.ini playbooks/rollback.yml
ansible-playbook --limit nodea.example.test -i inventories/lab.ini playbooks/rollback.yml

Rollback is an automation path with its own test, not a promise to edit hosts manually after failure. Preserve the previous artifact, limit the host pattern, execute serially where availability requires it and verify the restored service from the client side.

Knowledge checks with explained answers

What is the best first exam action?
Read every requirement and map it to hosts, files and acceptance evidence.
What should remain after completion?
A repeatable project whose second run is clean and whose services pass tests.
How should a failed task be approached?
Separate selection, connection, privilege, input, module and service layers.
What should the second run prove?
The desired state is already present, so state-based tasks normally report no changes.
Why use FQCN module names?
They identify the collection that owns the module and avoid ambiguous short names.
Is check mode proof of a safe change?
No. Module support varies and check mode cannot predict every external effect.
Why apply a host limit first?
It proves inventory selection and reduces the blast radius of a mistaken pattern.
What belongs in Git?
Reviewed automation source and non-secret inventory data, never vault passwords, private keys or tokens.

Guided lab and independent challenge

  1. Recreate the starting state on nodea and nodeb and record the inventory graph.
  2. Run syntax and check mode against nodea only; explain every predicted change.
  3. Run the play against both nodes and verify the result from the service or client side.
  4. Run it again and investigate any unexplained change.
  5. Inject the lesson-specific failure, capture the failed layer and execute the reviewed recovery.
  6. Complete the same end state from a fresh Git checkout without copying commands from the article.

Repeat the challenge against fresh nodes. The completed state, not a remembered command sequence, is the assessment target. Save syntax output, first and second recaps, a negative test and the rollback result.

Primary references

Advertisement