A quality gate should block a change only when the evidence indicates a defined level of product or delivery risk — not whenever any test, anywhere, fails to pass.
“All tests must pass” reads like a safety policy. In practice, most teams that adopt it as a literal rule eventually build an informal escape hatch around it — a re-run button, an admin override, a Slack thread where someone with the right permissions manually unblocks a release because the alternative is shipping nothing while an unrelated, already-tracked flake sits red. The rule didn’t make the system safer. It made the system dishonest about how release decisions actually get made, because the real decision moved out of the pipeline and into a private conversation.
A gate that blocks indiscriminately trains engineers to route around it. A quality gate should block a change only when the evidence indicates a defined level of product or delivery risk.
Why “all tests must pass” is not a complete policy
The rule conflates two very different situations: a deterministic test failing because it caught a real defect, and a test failing because the environment it depends on happens to be unavailable right now. Both produce the same red mark. Only one of them should stop a release. A policy that can’t tell them apart isn’t a policy — it’s a coin flip dressed up as rigor, and the workaround culture that grows around it is a rational response to that gap, not a discipline failure on the team’s part.
Required versus advisory evidence
Not every signal belongs at the same authority level. A useful gate design separates checks the team has agreed genuinely must pass — required evidence — from checks that inform a decision without automatically blocking it — advisory evidence. Collapsing that distinction is what forces the informal-override pattern into existence in the first place; a gate with no advisory tier has no honest way to say “we saw this, we’re proceeding anyway, and here’s why” without someone quietly bypassing the mechanism.
Risk-based blocking rules
Decision
Blocking should be a function of confirmed risk, not of any failure existing
A confirmed product-behavior failure on a required test blocks. An environment outage on a required test pauses and triggers environment recovery — it does not falsely indict the change. A known, accepted, low-risk issue on an advisory suite continues with the acceptance recorded, not silently ignored. A confirmed security exposure blocks regardless of anything else. A quarantined test failing again reports through its existing tracked issue and does not independently re-block the pipeline.
An example gate model makes the distinctions concrete:
Signal
Result
Action
Required deterministic test fails
Confirmed product behavior
Block
Required test fails
Environment unavailable
Pause and restore environment
Advisory suite fails
Known low-risk issue
Continue with recorded acceptance
Security control fails
Confirmed exposure
Block
Quarantined test fails
Existing tracked issue
Report; do not independently block
Human approval as a deliberate exception, not a default
Manual override has a legitimate place — but only when it’s a named, visible exception mechanism, not the pipeline’s actual failure mode operating in the shadows. A gate design that expects and accounts for human judgment on a defined subset of situations (a genuinely ambiguous failure, a business-driven exception under time pressure) is more honest than a rigid gate that gets quietly bypassed the same way every time something inconvenient happens.
Auditing false blocks and escaped defects
A gate’s design should be revisited using two failure modes, tracked separately: how often it blocks a release that later turns out to have been fine (a false block, which erodes trust in the gate and encourages workarounds), and how often a defect escapes past it into production (a missed block, which is the risk the gate exists to prevent). A gate tuned only against one of these numbers will quietly get worse at the other — a gate with zero false blocks that also lets defects through isn’t strict, it’s just badly calibrated in the other direction.
Key takeaways
Key takeaways
“All tests must pass” collapses two different problems — confirmed defects and environment noise — into one blocking rule, which is what creates informal override culture.
Separate required evidence from advisory evidence; both matter, but only one should automatically block.
Blocking should be a function of confirmed risk category, not of any failure existing anywhere in the pipeline.
Manual approval is legitimate as a named, recorded exception mechanism — not as the pipeline’s unacknowledged actual behavior.
Audit both false blocks and escaped defects. A gate is not strong because it blocks often; it’s strong because engineers trust the reason when it does.
Saved on this device —sign into sync your progress across devices.
✓ Marked as read
You've read 0 articles —sign into save your progress across devices.
Quarantine is a temporary risk-control mechanism for a specific, tracked problem — not a place inconvenient tests go to stop blocking the pipeline forever.
A red pipeline stage should contain enough structured evidence to identify the failed behavior, the affected system, a likely failure category, and the next investigation step — not just a final assertion.
Teams optimize CI for execution volume, pass rate, and runtime. None of those describe whether the pipeline actually helps anyone make a better engineering decision.