DevOps / Ansible Interview questions
How do you troubleshoot a failing Ansible task?
- Increase verbosity (
-v,-vvv, or-vvvv) to see the exact module arguments and, at higher levels, the raw SSH/connection details. - Check the error message and module return values - most modules return a descriptive
msgfield explaining exactly what went wrong. - Register the task's output and debug it to inspect intermediate values when a downstream task behaves unexpectedly.
- Run with --check --diff to see what change was expected, isolating whether the issue is in the logic or the actual execution.
- Test connectivity separately with an ad-hoc
ansible host -m pingto rule out SSH/auth issues before assuming the playbook logic is at fault. - Isolate with --limit and --tags to re-run just the failing host and task rather than the whole inventory and playbook.
ansible-playbook site.yml --limit web1.example.com --tags deploy -vvv
Most task failures fall into one of a few buckets - a connectivity/auth problem, a variable that resolved to something unexpected, or a module receiving arguments it doesn't accept in the current version - and narrowing which bucket applies with verbosity and targeted re-runs is usually faster than guessing from the top-level error alone.
More Related questions...