Menu

Measuring Whether Automation Helps

"Is this suite actually helping?" is a fair question that most teams cannot answer. A field note on instrumenting one suite for four weeks, and the metrics that changed the conversation.

5 min read

On this page
  1. The suite nobody trusted
  2. Measure behaviour, not size
  3. What the numbers said
  4. What changed
  5. The number I watch now
  6. Key takeaways

The suite nobody trusted

Last year I inherited a suite of around 1,400 checks guarding a payments product. The pipeline took seventy-one minutes on a good day. Roughly one run in three was retried at least once before someone believed the result, and “did it actually fail, or is it the suite?” was a standard stand-up question. At my first review the engineering manager asked the obvious question: is this thing actually helping us?

The honest answer was that nobody knew. The team could tell me how many checks existed, what the pass rate was, and how long the pipeline took. None of those numbers answers the question. A suite helps when people change decisions because of it: merge or hold, ship or roll back, investigate or ignore. So we spent four weeks measuring the suite’s behaviour instead of its size.

Measure behaviour, not size

We instrumented six metrics, all derivable from CI logs and the incident tracker, none requiring new tooling:

Metric Definition Healthy direction Warning sign
Feedback latency Commit pushed to trustworthy suite result Down; under ~15 minutes at the merge gate Slow creep upward, quarter over quarter
False-alarm rate Failures that did not indicate a product defect Below ~5% Anything sustained above 10%
Rerun rate Runs retried at least once to obtain a green Near zero Retries baked into the pipeline configuration
Signal response time Red suite to a human actively investigating Minutes Hours, or “we just restart it first”
Maintenance share Team effort spent keeping the suite alive Small, stable minority Growing quarter over quarter
Defect escape rate Production defects the suite could plausibly have caught Trending down Flat while the suite keeps growing

None of these is exotic. What surprised the team was that none of them had ever been written down.

Two practical notes on collecting them. First, classification beats precision: to compute the false-alarm rate we had an engineer tag every red run with one of five cause labels (product defect, environment, test data, check logic, unknown) which took under a minute per failure and was accurate enough to steer by. Second, the defect escape rate needs the phrase “could plausibly have caught” to stay honest. Not every production incident was ever automatable; counting only the plausible ones keeps the metric a measure of the suite rather than a measure of reality.

What the numbers said

Four weeks of data told a coherent story. Feedback latency averaged seventy-one minutes, which meant the merge gate could not serve the merge decision: developers opened the next piece of work rather than wait, and failures were context-switched into hours later. The false-alarm rate was 23%: nearly one failure in four was the suite crying wolf, mostly environment and shared test data. The rerun rate was 34%: retrying was not a workaround, it was the process.

The most damning number was signal response time: a red suite at 17:00 was typically first looked at the following morning. The suite was producing signal. The organisation had learned, rationally, to ignore it, because three years of false alarms had taught it to.

What changed

We deleted or quarantined 380 checks, most of them end-to-end duplicates of risks already covered lower down. Shared mutable fixtures were replaced with per-run data. Latency fell to fourteen minutes; false alarms fell under 4%. Only then did the interesting thing happen: response time collapsed on its own. When red became rare and usually real, people looked immediately, no process change required.

The maintenance share told the quieter story. Before the cleanup, one engineer was spending close to a third of their week feeding the suite; afterwards the same suite asked for a few hours. That reclaimed time funded the contract checks the team had been postponing for a year, the measurement work paid for the improvement work.

Defect escape rate was the last metric to move, and it moved only after response time did. That sequence taught me the value chain of a suite: detection enables response, response enables prevention, prevention is what shows up as fewer escapes. Teams that try to improve escapes directly, by adding more checks, are pushing on the wrong end of the chain.

The number I watch now

If I could keep only one metric, it would be rerun rate. It is a trust proxy the team computes every day without noticing: every retry is a small vote that the suite’s word is not enough. When that vote share drifts upward, everything else (latency, response, escape rate) is already on its way down. Trust is the asset; the other metrics are its shadow.

Key takeaways

  • A suite’s value is behavioural: people changing decisions because of it. Measure that, not its size or pass rate.
  • False-alarm rate and rerun rate are trust metrics. Once they degrade, every other number stops mattering.
  • Feedback latency is a strategy constraint, not an infrastructure detail: it decides which decisions the suite can serve.
  • Report trends, not snapshots. A suite improving on four metrics is an asset even while every absolute number is still mediocre.