Engineering
Test Flake Stabilizer
Find the real cause of flakes instead of wallpapering them with sleeps.
Use when
CI fails differently across comparable runs.
Cadence
When tests are inconsistent
Verification
The repaired test and full suite pass for the required consecutive-run streak.
Structured loop spec
| Field | Value |
|---|---|
| Name | Test Flake Stabilizer |
| Category | Engineering |
| Trigger | When tests are inconsistent |
| Objective | Find the real cause of flakes instead of wallpapering them with sleeps. |
| Allowed inputs | Relevant files, source notes, logs, tests, screenshots, metrics, or task state for this loop |
| Allowed actions | Define the exact scope, source of truth, and approval boundary.; Inspect current state and rank the highest-risk gap.; Make one small, reversible improvement.; Run the stated verification and record evidence.; Stop on success, budget, no progress, or approval required. |
| Verification | The repaired test and full suite pass for the required consecutive-run streak. |
| Stop condition | Stop when the verifier passes, the budget is exhausted, no progress is made, a blocker appears, or approval is required. |
| Budget | Set a time, turn, token, retry, file, or dollar cap before running the loop. |
| Approval boundary | Human approval required before publishing, sending, deleting, spending, changing accounts, touching production, or making reputational/legal/financial commitments. |
| Safe output | Pull request, patch, report, or evidence log |
| Works with | Claude Code, OpenAI Codex, Cursor, Gemini CLI, any tool-using coding agent |
Steps
- Define the exact scope, source of truth, and approval boundary.
- Inspect current state and rank the highest-risk gap.
- Make one small, reversible improvement.
- Run the stated verification and record evidence.
- Stop on success, budget, no progress, or approval required.
Prompt
Run the Test Flake Stabilizer loop. Use it when CI fails differently across comparable runs. Work in bounded iterations: inspect current state, choose the highest-risk gap, make one reversible improvement, verify it, and record evidence. Stop when The repaired test and full suite pass for the required consecutive-run streak. or when blocked, budget exhausted, or approval is required.Run in Claude Code
Paste this into Claude Code (or any tool-using agent) to run the loop bounded: one change per round, the same verification every round, durable state files, and explicit stop conditions.
Run the "Test Flake Stabilizer" loop from AI Loop Library (https://ailooplibrary.com/loops/test-flake-stabilizer/) as a bounded loop.
Goal: Find the real cause of flakes instead of wallpapering them with sleeps.
Rules: one change per round; run the same verification every round (The repaired test and full suite pass for the required consecutive-run streak.); append each round to docs/loops/test-flake-stabilizer/progress.md and update docs/loops/test-flake-stabilizer/state.json; stop on verifier pass, 8 rounds, 3 consecutive failed verifications, no progress, a blocker, or anything needing human approval (money, production, outbound, deletion). Finish with a proof report: rounds used, changes made, verification output, remaining risk, and the next human decision.