Identifying High-Value Automation
Not every check deserves to exist. A working method for ranking automation candidates by the decisions they serve, and for saying no to the rest.
6 min read
On this page
The backlog is a portfolio
Once a team accepts that automation should own only certain questions, the next problem arrives immediately: the candidate list is always longer than the capacity to own checks well. The temptation is to work through the backlog in the order it was written. That treats the suite like a queue of chores. It behaves much more like a portfolio: every candidate competes for the same finite attention, and each one carries both an expected return and an ongoing cost.
The return on a check is the value of the decision it accelerates: how often the question is asked, how bad a wrong answer is, and how much later the team would otherwise find out. The cost is everything the check will consume across its lifetime: authoring, environments, test data, maintenance, and the debugging time its future failures will demand.
Three questions that rank any candidate
Before writing a check, I ask three questions, in order:
- What decision does this serve, and how often? A check guarding the merge gate runs thousands of times a year; a check guarding an annual data-migration script runs once. Frequency multiplies both value and cost.
- What is the cost of being wrong, and how late would we otherwise learn? A defect caught in production after a week of customer impact justifies far more automation than one a reviewer would spot in a pull request.
- What will this check cost to own for a year? Include the flaky afternoons, the environment drift, and the test-data upkeep, not just the afternoon it takes to write.
Candidates that score high on all three are obvious. The interesting work is in the disagreements: checks with huge blast radius but unstable expectations, or cheap checks guarding journeys nobody would page anyone for.
Make the trade-off visible
A simple scorecard is enough. The point is not to manufacture a precise number; it is to force the trade-off into the open before a test title makes the answer feel inevitable. Rate the decision value and ownership cost as low, medium, or high, then record the lowest test level that can answer the question.
| Candidate | Decision served | Cost of being wrong | Ownership cost | Best home |
|---|---|---|---|---|
| Tax rounding for an order total | Merge every pricing change | High: money is wrong | Low | Unit |
| Payment request matches provider contract | Release a checkout change | High: payments fail at runtime | Medium | Contract |
| A signed-in user can update one address | Confidence in a critical journey | Medium | High: browser, account, data | End-to-end |
| Marketing copy appears in a banner | Visual review before merge | Low | High if made brittle | Manual review |
This makes the usual mistake easier to spot. Teams often rank candidates only by business importance, then give every important outcome an end-to-end check. That confuses the value of a decision with the price of one particular way of checking it. Tax rounding matters enormously; that is the reason to put it in a fast, cheap unit check, not to wrap it in a browser journey.
The converse is just as important. A low-frequency operation can deserve strong automation when the recovery cost is extreme. A yearly migration that can corrupt customer records is not made safe by its calendar position. Its runbook may need dry runs, idempotence checks, and an integration test against a production-shaped copy of the data. Frequency is a multiplier, not a veto.
Count the cost honestly
Build cost is visible in a ticket estimate; maintenance arrives later, in fragments, charged to whoever happens to be on call when the suite goes red.
The cost rises with everything a check needs that the product does not. Shared accounts, seeded data, clock control, third-party sandboxes, asynchronous queues, browser timing, and a particular ordering of tests are all small operational dependencies. None is automatically disqualifying. Each is a claim on future attention, so each belongs in the ranking.
A check is not valuable because it catches a bug. It is valuable when it catches an important bug early enough, reliably enough, and cheaply enough that people still believe it next year.
The practical question is not whether an end-to-end journey catches the defect. It usually does. The question is whether a lower-level check catches the same defect sooner, with a clearer failure, and without reserving another fragile seat in the suite. If it does, the higher-level check needs a different risk to justify itself.
Use thresholds, not a universal formula
Some teams turn the scorecard into a five-by-five matrix with weighted columns. That can help in a regulated or safety-critical domain, but it is usually more ceremony than the decision needs. A short set of thresholds travels better:
- Automate when the decision recurs, the failure is costly, and a stable signal is available at a proportionate level.
- Prefer the lowest level that observes the actual risk, not merely an adjacent symptom.
- Escalate to an end-to-end journey only when the risk crosses boundaries that lower checks cannot see.
- Stop when the maintenance estimate is larger than the decision value; write down the manual control that replaces it.
The last threshold is the one teams skip. “Do not automate” can sound like neglect, so a candidate stays on the roadmap until someone has a spare sprint. A recorded manual control is more honest: a release checklist, a peer review prompt, a sampled report, or an operational alert. It says how the risk is managed now and what evidence would justify revisiting the decision later.
The rejection list matters more
The highest-value artefact a strategy produces is not the suite. It is the written list of candidates the team has decided not to automate, with the reasoning attached. That list prevents the same argument returning every quarter with less context and more urgency.
One product team I worked with rejected a browser check for every combination of roles, locales, and feature flags. The proposed matrix would have created hundreds of journeys, most differing only in copy. We kept contract checks for permissions, unit checks for flag evaluation, and three end-to-end journeys for the business flows that actually crossed those boundaries. The rejection note named the gap: visual localisation regressions would be caught in review and by a small sampled smoke run. Six months later, when a locale-specific checkout defect appeared, the note made the next decision straightforward. The risk had changed; the portfolio changed with it.
That is what a useful backlog looks like. Not a pile of test ideas waiting for capacity, but a current set of bets: the risks being bought down, the checks earning their maintenance, and the tempting work that has been deliberately left out.
Key takeaways
- Rank automation candidates by the decision they accelerate, the cost and timing of a wrong answer, and the full-year cost of ownership.
- Business-critical behaviour does not automatically earn an end-to-end check; place it at the lowest level that can observe the real risk.
- Record the operating dependencies behind each candidate. Shared data, external systems, and timing are costs, not implementation details.
- Keep a rejection list with the manual control and revisit condition. Saying no with evidence is part of a test strategy.
Saved on this device —sign into sync your progress across devices.
✓ Marked as read
You've read 0 articles —sign into save your progress across devices.