Seven Automation Rules I Keep Reusing
The practical rules I use when automation has to survive retries, manual intervention, and changing infrastructure requirements.
Table of Contents
I used to judge automation by how much manual work it removed. That is useful, but incomplete.
The harder question is what happens six months later: requirements have changed, the original owner has moved on, a dependency is partially unavailable, and somebody needs to override the workflow during an incident. Automation that only works on its intended path is usually just a script with a schedule.
Across infrastructure and ML systems, I keep coming back to the same seven rules.
1. Give each workflow one job
A workflow should have a boundary that can be explained in one sentence. Reserve capacity. Validate a configuration. Apply an approved change. Reconcile observed state.
When one pipeline forecasts demand, allocates resources, writes configuration, and sends notifications, every retry becomes ambiguous. Which steps are safe to repeat? Which state is authoritative? Which failure should stop the rest?
I would rather connect several small workflows through explicit state than hide an entire operational process inside one large job.
2. Use the same execution path for humans and schedulers
Manual intervention is not an exception to the system. It is one of the system’s inputs.
If an operator button runs different code from the scheduled path, the least-tested path becomes the one used during an incident. I prefer one task handler with different triggers: schedule, API, command line, or an approved operator action.
That also makes overrides auditable. The system can record who requested the action without inventing a second way to perform it.
3. Make retries boring
Retries are normal in distributed systems. Timeouts do not tell us whether the remote operation failed; they only tell us that we did not receive an answer.
For that reason, externally visible operations need an idempotency key or a natural identity. Creating the same reservation twice, applying the same configuration twice, or charging for the same unit of work twice should not be possible just because a caller retried.
If I cannot explain what a second execution does, the operation is not ready for automation.
4. Assign one owner to each field
Many automation bugs are really ownership bugs. Two controllers both believe they can update the same value, and the result depends on timing.
For each important field, I want one authoritative writer under a clearly defined condition. Other systems may propose changes, but they should not silently compete to commit them.
This rule matters more than the choice of queue, database, or workflow engine. Clear ownership removes an entire class of races.
5. Separate intent from observed state
Requested state and actual state are different facts.
A deployment may be approved but not installed. Capacity may be reserved but not available. A configuration may be generated but rejected by validation. Collapsing these into one status makes the happy path look simple and every recovery path confusing.
I prefer a small state model that preserves both intent and observation. Reconciliation can then answer a concrete question: what must change to make reality match the approved intent?
6. Build the recovery path first
The success path is usually straightforward. The difficult design work starts with partial completion.
What happens if step three of five fails? Can the workflow continue after the dependency recovers? Does rollback make the situation safer, or would it destroy useful progress? Is a human allowed to take over without corrupting state?
I do not assume rollback is always the answer. In infrastructure work, forward recovery is often safer. The important part is that recovery is represented in the design, not left to whoever is on call.
7. Treat observability as part of the interface
An automation system needs to explain itself.
For any action, an operator should be able to answer: what triggered it, what inputs it used, what decision it made, what changed, and what it will do next. Logs alone rarely provide that view unless the state model was designed for it.
This is also why I avoid vague statuses such as failed when the system knows the difference between invalid input, unavailable capacity, rejected configuration, and a transient dependency error.
The standard I use now
I no longer call a workflow reliable because it completed a demo without human input. I look for something less exciting:
- a retry does not create damage;
- ownership is unambiguous;
- partial progress is visible;
- a human can intervene through the same controlled interface;
- and the next engineer can understand why the system acted.
Good automation is not the absence of people. It is a system that lets people spend their judgment where judgment is actually needed.