Introduction
Most engineering teams face the same issue at least once in their testing lifecycle: A test fails, testers rerun the pipeline, and the test passes. The test is the same, but it produces different results each time it is run. This is a kind of test that we call a flaky test.
There’s no definite answer to why a flaky test failed in the first place and passed the second time. But if that happens enough times, your test suite becomes clerical work instead of actual QA.
The usual solution is to delete the test or leave a comment, but neither of these gives you any coverage that you may need later. A better solution is to quarantine a flaky test, with some conditions attached. In this guide, we’ll learn what it means to quarantine a flaky test and how to do it.
What Are Flaky Tests
Flaky tests are the kind of tests that produce different results each time they run without any changes in the code. They may pass or fail inconsistently, which gives testers no clue about what is broken, if anything.
Flaky tests are a problem because they result in wasted time, wasted cost, and poor trust in releases—a flaky test can indicate that other tests that are actually failing might also be flaky, potentially resulting in inaccurate defect management.
What It Means to Quarantine a Flaky Test
Quarantining a flaky test means isolating an unreliable flaky test from your primary test suite so that its failures do not block continuous integration or deployment pipelines. Usually, a failed test blocks your deployment pipelines, which means the bug must be resolved and test must be passed before you can continue the integration. However, a flaky test is different from a failed test, so it requires quarantine.
Instead of outright deleting the test or ignoring its output, a quarantined test is moved to a separate, non-blocking execution lane, so the test keeps running but stops blocking deployment. When the test is quarantined, it can still run and appear in reporting similar to a normal test, but it doesn’t stop the integration.
Quarantine vs. Skip vs. Delete Tests
Quarantining a test is different from skipping or deleting it.
When you skip or delete a test, you can’t run it, its result won’t be recorded, it cannot block merges, and it provides no data for diagnosis.
However, when you quarantine a test, you can still run it, record its results, retain its coverage, and get the full data for diagnosis while continuing to merge. You can also remove a test from quarantine after the underlying problem is identified.
When Should You Quarantine a Test
Every quarantined test is an unresolved problem in your application, so the bar of uarantining test should be based on real issues in the test. Quarantine a test when:
The results are non-deterministic: Quarantine the test if the code remains the same, but the results are different. If it fails consistently, it is a bug report, not a quarantine case.
It has a measurable failure rate: the failure rate between 1% and 5% is a common threshold.
It has actually blocked someone: A test that actually blocks a pull request is more urgent than a test that is not actively blocking anything.
It is not covering something critical: A test that is covering something critical like payment processing, authentication, or data integrity cannot be “saved for later.” You have to fix critical tests urgently.
If your test suite has a lot of quarantined test cases (more than 2%), there might be a problem with your test architecture.
How to Quarantine Flaky Tests: A Step-by-Step Process
Here’s a step-by-step guide on how to quarantine a flaky test:
Step 1: Detect Flakiness Automatically
Manual flakiness detection does not scale. Here are two reliable ways to catch the flakiness automatically:
1. Repeat runs: Run the same test multiple times against the same commit. Playwright supports this with --repeat-each=5. Most test automation frameworks have an equivalent. Any test that produces mixed results across those runs is flaky by definition.
2. Historical tracking: Record pass and fail results for every test across every run, then calculate failure rate per test over a rolling window. A test failing 3 out of 100 runs on the same branch is flaky, and you now have a number to point at.
Step 2: Split Your Suite Into Blocking and Non-Blocking Stages
Your test suite and deployment pipeline need two lanes:
1. The blocking stage contains everything that must pass before a merge. This is your required check set. It should be fast, stable, and absolutely trusted. If something in here fails, work stops.
2. The non-blocking stage runs the quarantined tests. It executes on the same commits, produces the same reports, and fails in its own lane without touching merge status. Give the non-blocking stage its own dashboard. Teams that route quarantine results into the same view as everything else tend to lose track of them.
Step 3: Tag or Manifest the Quarantined Tests
You need a machine-readable record of what is quarantined and why. Two approaches are good here:
1. Tagging in code: Add an annotation to the test itself, with structured metadata in the body. See the example below.
@quarantine(
owner: "priya.n",
reason: "intermittent timeout on checkout step, ~6% fail rate",
ticket: "QA-1842",
expires: "2026-10-15"
)
As a result, the context lives next to the test, so anyone reading the file knows immediately.
2. A manifest file: Keep a single file, YAML or JSON, listing every quarantined test with the same fields. Your test runner reads it and routes accordingly. You get one place to look, and you can quarantine without touching test code.
Step 4: Assign an Owner and Open a Ticket
The owner is a person who is in charge of the test case. Assign the developer who owns the code under test, or who wrote the test, or who touched it last—the rule should be consistent.
Open a real ticket in the system your team actually uses, such as GitHub or any native defect tracker in your test management platform. The ticket should carry the failure rate, a link to a failing run, the suspected cause if anyone has a guess), and the expiry date.
Step 5: Set an Expiry and Enforce It
Every quarantine test entry should have a date. Two weeks is a reasonable default. Longer than a month can lead to delays, and the date should be enforced. Before the entry expires, you should either fix the test and graduate it back or renew the entry if you need more time. Renewals should be capped by a small number so the solution is prioritized.
What Is the Graveyard Anti-Pattern and How to Avoid It
The Graveyard Anti-Pattern occurs when flaky tests are moved into quarantine and then forgotten. Instead of serving as a temporary holding area while issues are resolved, the quarantine becomes a permanent resting place for neglected tests. Over time, test coverage silently degrades, and teams lose visibility into real failure signals.
How the Graveyard Anti-Pattern Develops
Common reasons behind the graveyard anti-pattern are:
1. Quick-fix mentality: Developers quarantine failing tests to unblock builds quickly without opening follow-up tracking tickets.
2. Lack of ownership: Quarantined tests lack assigned owners or clear expiration dates, leaving no one accountable for fixing them.
3. Out of sight, out of mind: Non-blocking execution results are ignored, hiding persistent failures and regressions until major outages occur.
How to Avoid the Graveyard Anti-Pattern
Here’s how to avoid the graveyard anti-pattern:
1. Enforce mandatory metadata: Require every quarantined test to specify an owner, an issue tracker ticket, a specific reason, and an expiration date.
2. Set strict quarantine limits: Cap the total number of quarantined tests (e.g., maximum 5% of the test suite). Require resolving existing quarantined tests before adding new ones once the cap is reached.
3. Automate expiration alerts: Trigger automated notifications or build warnings when a test exceeds its scheduled time in quarantine.
4. Conduct regular triage reviews: Review quarantined tests during weekly engineering syncs to ensure active investigation, graduation, or permanent deletion.
How to Graduate a Test Back Out of Quarantine
Getting a test out of quarantine should be as clearly defined as putting it in. Otherwise, tests either linger indefinitely or get rushed back into the main suite, only to start blocking builds and frustrating the team again.
To prevent premature graduation, establish a strict stability bar. A standard benchmark requires the test to pass 50 consecutive runs in the non-blocking execution lane without a single failure. For tests that were severely flaky, increase this threshold to 100 consecutive green runs before considering them stable.
Here’s how the sequence should go:
1. Fix the root cause, not the symptoms: Avoid quick fixes like adding retry wrappers or extending arbitrary sleep timeouts. Instead, replace static waits with dynamic, event-driven assertions, isolate test data using unique identifiers per test run, ensure proper setup and teardown of environment state, and mock or stub unstable external dependencies. Band-aid fixes merely conceal underlying instability, guaranteeing the test will flake again.
2. Let the fix soak in CI: Keep the test in the non-blocking quarantined lane while it accumulates test runs across various branches and builds. For example, if your CI pipeline executes 20 times per day, completing a 50-run stability requirement will take roughly two to three days. Resist the urge to shortcut this phase by running the test locally in a loop, as local environments rarely replicate the concurrency and network conditions of CI runners.
3. Verify stability against metrics: Review actual build history logs and telemetry rather than relying on gut feeling or memory to confirm that the stability threshold has been reached without intermittent failures.
4. Promote back to the blocking suite: Remove the test from the quarantine manifest or delete its code annotation, close the tracking ticket, and restore the test to the primary blocking stage where failures halt deployment pipelines.
5. Monitor closely post-graduation: Track the test’s performance during its first week back in the blocking suite. If it fails due to flakiness again, return it immediately to quarantine and mark it for rewrite or deletion, as failing multiple graduation attempts indicates fundamental design flaws.
A pro tip: Continuously track two key performance indicators: median quarantine duration and overall graduation rate. If median duration rises, expiration policies are not being enforced effectively. If the graduation rate drops below 50%, it indicates that most quarantined tests should be deleted rather than repaired, saving valuable engineering overhead.
TestFiesta Turns Flaky Test Chaos Into a Queue You Can Actually Clear
Managing flaky tests effectively requires robust tracking and accountability. TestFiesta simplifies this workflow by serving as a centralized platform for test results, historical metrics, ownership, and quarantine statuses.
Here is how TestFiesta streamlines flaky test management from detection to graduation:
- Automated Tracking & Flakiness Trends: Instead of parsing complex CI logs, TestFiesta automatically gathers failure rates over time and highlights flakiness trends across your runs.
- Clear Ownership & Expiration Tracking: Quarantined tests are assigned directly to owners and linked with strict expiration deadlines, preventing them from being forgotten in config files.
- Data-Driven Graduation: When a test is ready to return to the blocking suite, TestFiesta provides verified run history to confirm stability before graduation.
FAQs
Does quarantining a test slow down my CI pipeline?
Yes, quarantining a test can slightly slow down your CI pipeline because quarantined tests still run. The delay is usually brief because the non-blocking stage runs in parallel with everything else.
Can I automate the quarantine process entirely?
Not entirely, but you can automate the quarantine process largely. Detection, routing, and expiry reminders can all be automated. But the decision to quarantine a test and the assignment of an owner should stay manual.
What if a quarantined test is actually catching a real bug?
Quarantine tests can sometimes actually catch a real bug, and it’s the main risk of quarantining a test. Before quarantining, check whether the failure correlates with specific code changes rather than appearing at random. If the failure rate jumps after a deploy, treat it as a regression first and investigate before routing it to quarantine.








