Why automations break — and the four properties of ones that don't

Most automations fail the same way — silently, on a Tuesday — and the fix is four unglamorous properties, not a new tool.

// key takeaways
  • Automations break on the day a step fails, runs twice, or meets an input nobody imagined — not in testing.
  • The duct-tape lifecycle ends the same way every time: a person checking by hand the work the automation was built to remove.
  • Four properties separate a script from a system: retries with backoff, idempotency, an audit trail you can replay, and a place for humans.
  • No-code tools are fine for glue; graduate when money, compliance, or volume ride on the result.
  • An LLM step inherits every reliability problem of the automation around it.

Week one, the automation works: every web form becomes a CRM record and a welcome email. Month three, it fails silently on a Tuesday, and nobody notices until a customer does. Month six, nobody trusts it, and someone checks every run by hand — the exact job it was built to remove.

Automations break because they're built for the day everything goes right, with no plan for the day a step fails, runs twice, or meets an input nobody imagined. The tool is rarely the problem. Automations that last share four properties: retries with backoff, idempotency (safe to run twice), an audit trail you can replay, and a place for humans when reality goes off-script.

The duct-tape lifecycle

The duct-tape lifecycle is the path most quick automations follow: it works, then it fails silently, then nobody trusts it, then a person quietly shadows it by hand. Each stage feels reasonable from the inside, which is why a team can pay for an automation and the manual process at the same time without noticing.

It starts well: two tools connected in an afternoon, tested on three examples someone had on hand. Then the world moves. A service goes down for a few minutes, a webhook arrives twice, someone renames a form field. The error lands in a run history nobody opens, or the run stops halfway: CRM updated, email unsent.

The first person to notice is a customer. After the second incident, the team stops believing the output, and someone starts a spreadsheet "just in case," checking each run and redoing steps by hand. Now the business pays for the automation and the manual process — plus the confusion about which one is the source of truth.

That's why "it worked in testing" means so little. Testing covers the inputs you imagined; production is made of the ones you didn't. An automation tested only on the happy path, where everything goes right, has been tested on the part that was never going to break.

Rule of thumb: An automation you have to check by hand isn't an automation. It's a second job with a worse interface.

What actually breaks (it's rarely the tool)

What breaks is almost never the automation tool itself. It's the world around it: APIs that time out, webhooks that fire twice, runs that die halfway, fields renamed upstream, rate limits, and the one person who understood the setup leaving. Each failure has a recognizable symptom and a property that prevents it.

Ask why automations break and you'll hear tool names. Here's what we find.

What failsWhat you seeThe property that prevents it
An API times outA record is simply missingRetries with backoff
A webhook fires twiceTwo orders, two welcome emailsIdempotency
A run dies at step 3 of 5CRM updated, email never sentIdempotency, then replay
A field is renamed upstreamBlank values, silentlyA place for humans
A rate limit hits at peakBursts of failures at the busiest hourRetries with backoff
The person who knew it leftNobody can say what it doesAn audit trail you can replay

None of the fixes is "switch tools." Every row is a design decision, which is why moving a fragile automation to another platform usually reproduces the same failures on a new subscription. The rows also interact: a retry after a timeout is exactly how a double run happens. The properties work as a set.

Schema drift, the renamed field, is the quietest failure: someone renames "Phone" to "Mobile" in the form builder, and the automation keeps running, faithfully copying blanks. Only validation that routes unrecognized input to a person catches it.

The four properties of an automation you can trust

Strip the lists of workflow automation best practices down to what prevents real failures, and four properties remain: retries with backoff, idempotency, an audit trail you can replay, and a place for humans. Each has a plain definition, a failure it prevents, and a question you can ask of any automation you rely on today.

We call them the four properties and treat them as the definition of done.

Property 1: retries with backoff

Retries with backoff means that when a step fails for a temporary reason — a timeout, a rate limit, a service briefly down — the automation waits and tries again, waiting longer each time, before giving up and raising its hand. It prevents the most common silent failure: a brief blip that permanently drops a record.

Good automation error handling starts with one question: is this failure temporary or permanent? A timeout is temporary; the same request will probably succeed in a minute. An invalid email address or a missing required field is permanent; it will fail the same way on the fifth attempt.

Retrying blindly is worse than not retrying. A loop on a permanent error burns through rate limits, buries the real problem in noise, and, if the step isn't safe to run twice, repeats whatever half-succeeded. Hammering an overloaded service also prolongs its outage. So the waits grow, with a little randomness so stalled runs don't retry in unison; attempts are capped; the final failure is loud.

How to check: if the system your automation writes to is down for ten minutes, what happens to everything that arrives meanwhile? "It's lost" and "I don't know" are the same answer.

Property 2: idempotency — safe to run twice

Idempotency means running a step twice has the same effect as running it once. In business terms: the same order must not be created twice, and the same customer must not get the same email twice — no matter how many times the trigger fires or a retry kicks in.

Things run twice more often than people expect. Many services deliver webhooks "at least once," resending when they don't get a quick acknowledgment. People double-click. And a retry after a timeout is a double run whenever the first attempt succeeded and only the reply got lost.

Some steps are idempotent operations by nature, like an elevator call button: press it five times and one elevator comes. Setting an order's status to "shipped" twice leaves it shipped. Other steps are not: create a contact, add a line item, send an email. The design work is turning the second kind into the first.

The standard technique is an idempotency key: a unique ID for each piece of work, such as the form submission's ID or the order number, that travels with every attempt. Before acting, the step checks whether that key was already handled and, if so, returns the earlier result. Some APIs accept a key directly; otherwise the automation keeps its own record, or the database refuses the duplicate with a uniqueness rule.

How to check: in a test environment, send the same trigger twice on purpose. If you get two of anything, the property is missing.

Property 3: an audit trail you can replay

An automation audit trail records, for every run, what happened, why, and with which input: the trigger, the data each step received, what each step did, and where it stopped. Replay means that after you fix a bug, you re-run the failed cases from their original inputs instead of reconstructing them from memory and screenshots.

A log says "error at 10:42." An audit trail says which form submission started the run, that step three failed because the customer record had no delivery address, and what the input looked like. One tells you something went wrong; the other tells you what to fix and which runs to push through again.

Replay is only safe because of property two: re-running a half-finished run must not duplicate the half that worked. The trail also covers the last row of the table: the system explains itself after its builder leaves. In the document intake pipelines we build, every extraction is logged and traceable to its source document.

One caution: an audit trail that stores inputs is storing customer data. Keep it inside your own environment, record only the fields you need to replay, mask what's sensitive, and set a retention period — long enough to debug and replay, not forever.

Property 4: a place for humans

A place for humans means the automation has a defined destination for the cases it can't or shouldn't handle: an exception queue with a named owner, where each item arrives with its context and a clear next action. Without one, exceptions still happen. They just land in someone's inbox — or nowhere.

Field note: A regional e-commerce retailer ran storefront, delivery scheduling, and carrier coordination in disconnected tools, with exceptions — a wrong address, a weather hold — handled by phone and spreadsheet. The system we built is deterministic first: queues, retries, an audit trail, with exceptions triaged and handed to a person, context attached. The ops team runs the day from one screen, and exceptions get handled before the truck leaves instead of after.

Explicit states are the other half. For a residential roofing company, every lead lands in one queue, and follow-up continues until the job is scheduled or closed — the only two exits. That's why no lead falls through the cracks: "no state" isn't an option.

A digital agency's reporting pipeline applies the same idea to data: scheduled transforms with checks between stages, so a failed stage stops and tells someone instead of producing a confident report from half the numbers.

How to check: when this automation meets something it can't handle, whose screen does it land on, and can they resume the run from where it stopped? If the answer is a shared inbox nobody owns, the property is missing.

No-code is fine — until it isn't

No-code tools like Zapier, Make, and n8n are the right choice for notifications and low-risk glue between apps: fast to build, cheap to run, easy to change. Graduate to custom code when money, compliance, or volume ride on the result, when the branching logic explodes, or when you need to replay failures reliably.

We build custom systems, and we'd still tell you that plenty of automations should stay where they are. The platforms aren't the weak point; all three offer retries, error handling, and run history in some form. If you're shopping for Zapier alternatives because a workflow keeps failing, check the design first: a design problem moves with you, and most Zapier reliability complaints we hear are design complaints. The n8n vs Zapier choice is about fit, not reliability — n8n can be self-hosted and handles heavier logic; Zapier connects a vast catalog of apps with little setup.

No-code (Zapier, Make, n8n)Custom code
Best forNotifications, low-risk glueMoney, compliance, volume
Branching logicFine until it's a mazeStays readable, with tests
Duplicate protectionPossible with lookups; easy to missDesigned in: keys, uniqueness rules
Audit trail and replayRun history; depth variesAs detailed as the work needs
Who can change itAnyone with accessAn engineer, with review

Graduating rarely means rewriting everything. Often the right move is to keep the glue and move only the step that carries the risk, such as creating the order, into code with the four properties built in. That's how we approach business automation, with the integrations underneath handled as software and data engineering.

When not to do this: If an automation posts to a team channel when a form arrives, and a missed message costs nothing, leave it alone. Don't rebuild it — not with us, not with anyone. Rebuild when you can name what a failure costs.

Agents don't fix this — they inherit it

An AI agent is still an automation: a workflow that calls a language model at some of its steps. It inherits every failure mode of a regular automation and adds one, a step that can be confidently wrong. An LLM step needs the same four properties, plus checkpoints where a person approves anything irreversible, expensive, or customer-facing.

The pitch often runs the other way: the agent is smart, so it will handle the errors. It won't handle a timeout better than a retry policy, or notice a duplicate email unless the step is idempotent. Its own failure mode, output that looks right and isn't, calls for validation after every model step and a record of what the model saw and said. We build agentic systems, and the ones worth trusting look exactly like this. More in what an AI agent actually is: a workflow with checkpoints.

Run it yourself

For each automation your business relies on, answer four questions in writing: what happens when a step fails, what happens if it runs twice, can you see what happened, and who gets the exception. Score one point per clear answer. Anything scoring under three that touches money gets rebuilt first.

List every automation someone would notice within a day if it stopped. For each, write down:

  1. On failure: what happens when a step fails? (Retries with backoff.)
  2. On a double run: what happens if the trigger fires twice? (Idempotency.)
  3. Visibility: can someone other than the builder see what a run received and did? (Audit trail.)
  4. Exceptions: who gets the case the automation can't handle? (A place for humans.)

Score a point only when the answer is specific. "It retries three times, then posts to the ops queue" earns the point. "I think it emails someone" doesn't. A one or a two is the duct-tape lifecycle, waiting for its Tuesday.

Anything under three that touches money — orders, quotes, anything with a price on it — gets rebuilt. The rest can usually be fixed in place: turn on the platform's retries and error alerts, add a lookup before each create step, name an owner for failures. A four in a no-code tool is a four; leave it alone.

If the rebuild list outruns your capacity, rank it by hours times the cost of a mistake — the scoring behind how we decide what to automate first, which comes with the automation triage worksheet.

An automation earns trust the way a colleague does: by being predictable on the bad days, not just the good ones. The four properties make that a design decision instead of a hope. None of it is clever, and that's the point: the part of a system that keeps it trustworthy should be boring enough that nobody talks about it, least of all the person who used to check it by hand.

Questions we get asked

Why do automations fail?

Most automations fail because they were built for the day everything goes right: the inputs someone tested with, on a day every connected service was up. They break when a step times out, a webhook fires twice, a run dies halfway, a field gets renamed upstream, or a rate limit hits at the busiest hour. The tool is rarely the cause; the missing plan for failure is.

What is idempotency in automation, in plain terms?

Idempotency means running a step twice has the same effect as running it once. The same order is not created twice and the same email is not sent twice, even if a trigger fires again or a retry repeats work that already succeeded. It is usually built with an idempotency key: a unique ID for each piece of work that the system checks before acting.

When should a business move off Zapier or Make?

Move a workflow off Zapier or Make when money, compliance, or real volume ride on its result; when its branching logic has grown into a maze nobody can follow; or when you need to replay failed runs reliably after a fix. Notifications and low-risk glue can stay where they are. Often the right move is partial: keep the glue and rebuild only the step that carries the risk.

What is an audit trail in workflow automation?

An audit trail is a record of every automation run: what triggered it, the input each step received, what each step did, and where and why it stopped. A good one lets someone other than the builder answer 'did this happen?' in minutes, and lets you replay failed runs from their original inputs after a fix. Because it stores customer data, it needs masking and a retention period.

How do you handle errors in workflow automation?

Sort errors into two kinds. Temporary errors, such as timeouts, rate limits, or a service briefly down, get retried automatically, with longer waits between attempts and a cap on how many. Permanent errors, such as invalid data, a missing record, or a renamed field, are not retried; they go to an exception queue with an owner, the failing input, and a clear next action. Every step should be safe to run twice.

Are AI agents more reliable than regular automations?

No. An AI agent is an automation with language-model steps, so it inherits every reliability problem of a regular automation, including timeouts, duplicate runs, and partial failures, and adds one more: output that looks right and isn't. Agents need the same retries, idempotency, audit trail, and exception handling, plus validation after each model step and a human checkpoint before anything irreversible or customer-facing.

// written by

Stefan Jovanovic · Partner · Frontier & Agents

Stefan lives on the frontier: testing new models the week they ship, and building multi-agent systems — swarms of specialized agents that carry real workflows end-to-end.

the people behind Tensorika →

// talk to an engineer

Somewhere in your business, a workflow is eating hours it shouldn't.

Send us the process. We'll score it, tell you what to automate first, and build the version that still works next year.