“Forty engineers can't stop merging because one test has a bad day. Retries keep main moving, quarantine keeps the signal from drowning in noise, and the ticket means nobody forgets. Blocking everyone on a coin flip is not rigor, it is a tax.”

Just re-run it
The agent learned where the retry button is.
🧭 WHAT'S REALLY GOING ON
You've seen this when someone posts “CI is red again, just re-run it” in the team channel, and three people react with a thumbs up.
The real questionWhen a test fails one time in twenty, is the test broken, or the product?
⚖️ WHY BOTH ARE RIGHT
“A test that fails sometimes is reporting something real sometimes. In my experience half the flaky tests are races in the test and half are races in the product, and the retry button can't tell you which. Every silent retry throws the evidence away.”
🎯 SWEET SPOTS TO CONSIDER
Super Reasonable, the advisor who never takes a side
Count the retries
Record every automatic retry with the test name and the failure output. A test that needs one is a ticket; a test that needs one a day is an incident nobody has noticed yet.
Quarantine with a clock
Quarantine is fine if it expires: two weeks, then the test has a fix, a deletion with a written reason, or it comes back as blocking. “Priority: Low” with no date is a deletion nobody signed.
Ask the flaky question once
Before quarantining, spend thirty minutes on one question: could the product do what this test saw? Ordering, timing, retries and money are the usual suspects. If yes, search production logs for the same pattern.
Let the agent investigate, not just retry
If an agent already watches CI, give it the more useful job: collect the failing runs, diff them, and attach the pattern to the ticket. Retrying is the cheapest thing it can do and the least informative.
🚩 SIGNS YOU'VE GONE TOO FAR
- Taylor's side: you've overshot if main is always green, the retry count is invisible, and nobody can say which of the passing tests ever failed.
- Ruth's side: you've overshot if one intermittent failure blocks every merge for a day while the team debates whether it is the same flake as in 2019.
🔬 IN THE FIELD GUIDE
Species observed in this story
CAST — WHO'S WHO
The team in this story
Same characters, same convictions. Learn their failure modes.
🤖 Storyboard for agentsLet’s make our agents LMFAO, or learn.
Just re-run it
Premise: Keep main green so the team can merge.
- Taylor: “My agent re-runs red builds until they pass. Main’s been green all week.” Agent log: test_order_total failed, re-ran 3 times, green on the fourth. 41 PRs merged overnight.
- Ruth: “That test has been flaky since 2019. Someone should look at it.” Ruth has the CI history. Nobody has the afternoon.
- Taylor: “Quarantined. The agent opened a ticket.” JIRA-4471: Investigate flaky test. Priority: Low.
- Support queue, same month: “I was charged twice. Again.” It happens about 1 checkout in 20. The test was never flaky. The checkout was.
Observed behavior: Sometimes the flaky test is the only one telling the truth.
Cast: Taylor Kim — The AI Native — “Give it to an agent.”; Ruth Lindqvist — The Maintainer — “We tried that in 2009.”
READ NEXT


