Skip to main content
Back to insights

Make Integration Retries Safe With Idempotency

Networks fail after a request has been accepted but before the caller sees the response. Retrying without a stable operation identity can create duplicate orders, payments or updates.

  • Architecture
  • Integration

Arinao Tshamano8 October 20262 min read

Networks fail after a request has been accepted but before the caller sees the response. Retrying without a stable operation identity can create duplicate orders, payments or updates. A timeout does not prove that an order request failed; the first attempt may have completed before the response was lost. A blind retry can create duplicate work. Give the operation a stable key so the receiver can recognise a repeated request and return the recorded outcome.

What good engineering looks like

How will the receiving system recognise that the same business operation has arrived again? Answer this for a named workflow and an accountable team. Identify the information needed at the moment of decision, the system that can settle a dispute, and the route for correcting dependent copies. The technical design should make those business rules visible to the people who operate and support the handoff.

  • Assign a stable operation key before the first attempt.

  • Store the outcome against that key at the system that performs the mutation.

  • Return the prior result for a repeat request and test timeout, replay and conflicting payload cases.

  • A named owner for a rejected, delayed or repeated item.

  • Evidence that the receiving system applied the intended business state.

A practical starting point

  1. Choose one real case and write down its trigger, expected outcome and responsible team.

  2. Simulate a timeout after the receiver commits, then retry with the same key.

  3. Confirm that the order or payment-like action is recorded once and that the caller gets a usable result. Test a different payload with the same key as an error rather than silently treating it as the original request.

  4. Record what the test exposed, who will resolve each open question and how the fix will be checked.

Keep the first design small enough to review with the people who run the process. Test an exception alongside the normal case. A successful transport test shows that data moved; it does not by itself prove that the receiving team can make the right decision or recover from a partial failure.

The decision to make

Idempotency prevents one class of duplicate mutation. It does not replace reconciliation, validation or a clear retention policy for keys. Decide what level of timing, traceability and recovery this workflow actually needs. Document assumptions that remain untested and revisit the choice when a partner, process or business rule changes. The point is a dependable operating decision, not simply a working interface.

Explore Algoza's related services: https://algoza.co.za/services/integration

Apply the thinking

Working through a related technology decision?

Share the operational context, current systems, constraints, and decision you need to make.

Discuss a requirement