Draft / owner approval pending
The retry that did the job twice
A timeout does not tell you whether the remote action failed. Good automation needs a way to ask what already happened.
Prepared for Sonny Saggar’s review. Not yet approved or published.
The screen says the request timed out. The operator clicks again. Two confirmation emails arrive.
This is a hypothetical failure, but it describes a common design problem: the caller did not receive a response, while the receiving system may already have completed the action. Repeating the request can turn uncertainty into a duplicate.
Treat uncertainty as a state
A useful automation distinguishes not attempted, accepted, completed, failed, and outcome unknown. The last state is inconvenient. It is also more honest than calling every missing response a failure.
Google's Site Reliability Engineering discussion of overload describes how retries can contribute to cascading load and recommends bounded approaches rather than uncontrolled retrying. That is an operational reason to design retry behavior deliberately. It does not mean that backoff alone prevents duplicate business actions. [1]
The business question is separate: can the recipient recognize the same logical operation when it arrives again? A delay between attempts changes timing. It does not supply identity.
Give the job an identity
Here is a proposed acceptance contract for a document-delivery automation. Create an operation identifier before the first attempt. Preserve the intended recipient and content version with it. Reuse that identifier for retries of the same operation. If the content or recipient changes, make the change explicit rather than silently reusing the old identity.
The receiving side needs durable evidence of what it accepted or completed. A local flag that disappears when the process restarts is not enough for a durable guarantee. The precise mechanism depends on the service, and an external provider's documented behavior must be checked rather than assumed.
Test the point between success and confirmation
The revealing test is to interrupt the workflow after the remote action succeeds but before the caller records the response. Restart it. Then inspect the actual outcome, not just the local log.
Also test conflicting retries, expiry, and a provider that cannot report status. The correct fallback may be a review task with the original evidence, not another automatic attempt. A person asked to investigate should see the operation identifier, timestamps, and the last known provider response.
These are design proposals for a bounded workflow, not proof that any current integration delivers exactly once. The claim must be earned against the actual provider and persistence layer.
The promise of automation is not that nobody ever looks at a failed job. It is that the machine preserves enough state to avoid repeating preventable work and gives the person a precise question when uncertainty remains. Before removing the retry button, make sure the system knows what a retry means.
Sources and scope
- Google SRE: Handling Overload
Retry behavior, load amplification and bounded retries; not exactly-once business-action guarantees. Checked 2026-09-25.