01Hypothesis
Running every playbook twice in CI and failing the build on any reported change in the second run will surface non-idempotent tasks that reviewers miss.
02Method
Added a second consecutive playbook run to a scratch pipeline, parsing the recap for a non-zero changed count and failing on it.
03Findings
The idea works and it found real non-idempotent tasks quickly, mostly command and shell tasks without a creates or changed_when guard.
Parked rather than concluded because the false positive rate is annoying. Some tasks legitimately report changed on every run, service restarts triggered by handlers being the common case, and separating legitimate from accidental needs a per-task allowlist. That allowlist is maintenance, and I have not decided whether it costs less than the reviews it replaces.
Would restart this if I were maintaining a larger playbook set.