Azure Functions retries: make repeated events safe for your application
Design Azure Functions handlers for repeated delivery by matching trigger retry behavior to durable idempotency, transaction boundaries and failure recovery.
11 Oct 2026, 08:39 UTC

Read the retry behavior of the actual trigger
Azure Functions can receive work from HTTP requests, queues, event streams, timers and other integrations. These triggers do not all share one retry policy. Some behaviors belong to the Functions runtime, some to the extension, and others to the underlying service. Identify the trigger and extension version before deciding how an event will be retried after a failure.
A retry repeats execution after an error, but it cannot know whether an earlier attempt already changed an external system. A timeout might occur after the database committed or after a remote API accepted a request. The application therefore needs a way to recognize repeated work independently of whether the previous invocation appeared successful to its caller.
Choose a durable operation identity
Use a stable identifier for the business operation, such as a payment event ID or an order transition ID. An invocation ID identifies one execution and is usually unsuitable as the deduplication key for a redelivered business event. Store the operation key in durable state with a uniqueness constraint so concurrent attempts cannot both create the same result.
When the operation changes one database, write the result and its idempotency record in the same transaction. A check followed by an unrelated write leaves a race between two invocations. The database constraint or transaction should decide which attempt owns the change; an in-memory cache alone does not survive another instance or a restart.
Separate durable state from external effects
- Validate the event and derive a stable business-operation key.
- Claim or apply the operation through a transaction and uniqueness constraint.
- Persist any required downstream work in a durable outbox.
- Execute external effects with their own idempotency key where the provider supports it.
- Record completion and retain a recovery path for partially finished work.
A durable outbox helps close the gap between committing local state and requesting an external effect. Another worker can deliver the recorded task after a process interruption. It still needs bounded retries and a reconciliation strategy for an ambiguous remote response. Do not claim that an outbox by itself makes every distributed effect happen exactly once.
Classify failures before retrying them
Transient transport failures and temporary service throttling can benefit from retry with backoff. Invalid input, an unavailable permission or a permanently missing record needs a different response. Repeating those cases without a clear terminal or deferred state wastes capacity and obscures the underlying issue. Follow the trigger's documented failure handling and poison or dead-letter path where one exists.
Record the business key, stage, safe error category and attempt outcome. Avoid logging complete event payloads when they contain personal data or credentials. An operator should be able to find unfinished work and explain whether it is retryable, waiting on an external change or already applied.
Rehearse the ambiguous cases
Test duplicate delivery, two simultaneous executions and a crash immediately after the local transaction commits. Also test an external operation that succeeds while its response is lost. The handler is ready when these scenarios preserve the intended business state and expose a useful recovery record. Successful execution of one clean event is only the beginning of that verification.
References
- Azure Functions Error Handling and Retry Guidance — Microsoft Learn
- Designing Azure Functions for identical input — Microsoft Learn
Sources & further reading
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.