Menu

Making Pipeline Failures Explainable

A red pipeline stage should contain enough structured evidence to identify the failed behavior, the affected system, a likely failure category, and the next investigation step — not just a final assertion.

4 min read

On this page
  1. Why a red stage is not a diagnosis
  2. Minimum failure evidence
  3. Use failure categories, not adjectives
  4. Separate evidence from interpretation
  5. Improve the summary iteratively
  6. Key takeaways

A stage goes red and the report says: Expected true, received false. Nothing else. Whoever opens it next has to reconstruct, from scratch, what the test was doing, which system it was exercising, and whether the failure is worth ten minutes or two hours of their day — using only the last line the test happened to print before it stopped. That reconstruction work is not a debugging skill problem. It’s a design gap in what the pipeline chose to preserve at the moment of failure.

A final assertion tells you execution stopped. It rarely tells you why, and it almost never tells you what to do next. A pipeline failure should contain enough structured evidence to identify the failed behavior, affected system, likely failure category, and next investigation step.

Why a red stage is not a diagnosis

The last assertion a test executes is, definitionally, the last visible symptom — not the cause. By the time an order fails to reach CONFIRMED, the actual problem might be three services upstream, and the assertion has no way of knowing that. Treating the assertion message as the whole story is how ten-minute investigations turn into ninety-minute ones, repeated by a different engineer every time the same class of failure recurs.

Minimum failure evidence

A failed execution that’s actually useful preserves a specific, bounded set of context — not everything the framework could possibly capture, but enough to answer the questions an investigator will ask first:

  • Test identity
  • Commit and build identifiers
  • Environment and service versions
  • Input or fixture identifiers
  • Relevant logs
  • Network or API traces where appropriate
  • Screenshots only when visually meaningful
  • Timing information and retry history
  • Correlation identifiers linking this run to distributed traces
failure-evidence.jsonjson
{
"testId": "checkout.confirmed-order",
"commit": "a73d6f2",
"environment": "staging-eu-2",
"attempt": 1,
"durationMs": 20140,
"correlationId": "trace-6a41bd",
"services": {
  "checkout": "2026.07.23.4",
  "payments": "2026.07.22.9"
},
"artifacts": ["request-response.json", "service-events.json"]
}

None of these fields are exotic. Most systems already generate this data somewhere; the gap is usually that it’s discarded at the end of a passing run instead of attached to the failure that needed it.

Use failure categories, not adjectives

“Flaky” describes a feeling, not a cause. A short, fixed set of categories — product behavior, test implementation, test data, environment, dependency, capacity, timing or concurrency, unknown — gives every failure a first, provisional classification that narrows where to look, even before a human investigates further.

Separate evidence from interpretation

Raw artifacts should stay available in full, but the summary a person reads first should guide them, not bury them in logs:

This structure — observed, expected, correlated signal, likely category — does most of the triage work before anyone opens a debugger. It doesn’t claim certainty; it claims a starting point, which is the entire value proposition of a good failure summary.

Decision

Provide ownership context without hard-coding names

Map components and services to owning teams in configuration, not in the test itself, so a failure summary can say “route to checkout-quality” without anyone maintaining a list of individuals who happened to be on the team when the test was written.

Improve the summary iteratively

The set of fields worth capturing is never final. Every recurring manual investigation — the kind where someone has to go dig up the same piece of context a third time this month — is a signal that the failure summary is missing something specific, and a concrete instruction to add it. Treating failure evidence as a fixed schema, set once and never revisited, guarantees it slowly falls behind the system it’s describing.

Key takeaways

Key takeaways

  • Failure output should support the next action, not merely confirm that a test stopped running.
  • A fixed set of structured fields — identity, versions, correlation IDs, timing, retry history — turns investigation from reconstruction into lookup.
  • Fixed failure categories give every red result a provisional classification without pretending to establish certainty.
  • Separate the human-readable summary (observed, expected, correlated signal, likely category) from the raw artifacts backing it.
  • Every repeated manual investigation is an opportunity to add the missing signal to future failure reports.

Reliability6 min read

Diagnosing Flaky Tests: read the related article

Non-deterministic failures are defect reports with bad formatting. A field workflow for reproducing, triaging, and resolving flaky tests before the team stops believing the suite.