Courier API Failure Handling: Retries, Backoff and Idempotency

Any developer can call a courier API successfully once. The engineering is in the other cases: the timeout, the rate limit, the malformed response, the success that turns out not to have been one.

This article covers the patterns that separate an integration that holds up from one that quietly corrupts your data.

The problem in one sentence

A failed request does not mean the action did not happen.

If you send a shipment creation request and the connection times out, you know only that you did not receive a response. The shipment may have been created perfectly, with the carrier's acknowledgement lost on the way back. If you retry blindly, you have now booked the parcel twice, paid twice, and generated two AWBs for one box.

Almost everything below exists to deal with that single ambiguity.

Idempotency: the essential one

An idempotent request can be sent any number of times and have the same effect as sending it once. You achieve this by attaching a unique key to the request — typically derived from your own order or shipment identifier — which the carrier uses to recognise a repeat.

If the carrier supports idempotency keys, use them. If they do not, you have to build the equivalent yourself:

  • Record your intent to book before making the call, with a pending state
  • On an ambiguous failure, do not retry blindly — query the carrier first to see whether the shipment exists
  • Only book again if the query confirms it does not

This is more work than a retry loop. It is also the difference between a duplicate rate of zero and a duplicate rate that shows up on an invoice.

Retries with exponential backoff

When a call fails transiently, retrying immediately is the worst option: if the carrier is struggling, you are adding load to a system that is already failing, and you will be rate-limited for it.

Exponential backoff means waiting progressively longer between attempts — one second, then two, then four, then eight — with a cap and a maximum attempt count. Add a small random jitter, or every client of that API that failed at the same moment will retry at the same moment and recreate the spike.

Retry only what is worth retrying

This distinction is routinely missed. Retrying a permanent failure is pure waste.

  • Retry — timeouts, connection failures, 429 rate limits, 502/503/504
  • Do not retry — 400 bad request, 401 unauthorised, 404, validation errors

A 400 will be a 400 every time. It needs a human or a correction, not another attempt. If a 429 response includes a Retry-After header, honour it rather than using your own schedule — the carrier is telling you exactly what they will accept.

Queue rather than fail

When retries are exhausted, the work should go to a queue, not disappear. A queued booking is a delay; a lost booking is a parcel nobody ships and a customer nobody told.

A usable queue needs visibility — somebody has to be able to see what is stuck and why — and a way to replay items once the carrier recovers. A dead-letter queue nobody ever looks at is only marginally better than no queue at all.

Circuit breakers

If a carrier has been failing consistently for several minutes, continuing to send them traffic helps nobody. A circuit breaker watches the failure rate and, past a threshold, stops attempting for a cooling-off period, then lets a single trial request through to test recovery.

The operational benefit is that your own system stays responsive. Without one, every request queues up waiting on a timeout, and one failing third party degrades your entire application.

Timeouts

Set them explicitly, and set them shorter than you think. A default timeout of thirty seconds or more means one slow carrier can tie up your worker threads until the whole system stops responding. Separate connect and read timeouts, tuned to what the endpoint normally does, are part of the design rather than a detail.

Logging that is actually useful afterwards

When a carrier says "we never received that request", you need to be able to answer precisely. That means recording, for every call:

  • Request and response payloads, with credentials and personal data masked
  • Timestamps and duration
  • Attempt number and final outcome
  • A correlation ID linking every attempt for one shipment

That correlation ID is what turns an incident from an argument into a five-minute lookup.

Alerting on the right things

Alerting on individual failures produces noise that people learn to ignore. Alert on patterns instead: failure rate above a threshold, queue depth growing, a carrier returning an unrecognised status, or no successful call to a partner in an unexpectedly long period.

That last one catches the quietest and worst failure mode — an integration that stopped working entirely and is producing no errors because it is producing nothing at all.

Where we come in

Retries, idempotency, queuing and monitoring are part of how we build courier API integrations rather than hardening added afterwards. The same discipline applies across the other integrations we build.

More reading

How Courier API Integration Works

What actually happens when your software talks to a courier's API — the endpoints involved, the order they are called in, and the parts that are harder than they look.

Webhooks vs Polling for Shipment Tracking

Two ways to find out that a parcel moved. One is more efficient, one is more reliable, and most real systems end up using both — for good reasons.

Tracking Status Normalisation Across Carriers

Every carrier describes the same journey differently. Translating them into one consistent model is the largest hidden cost in multi-carrier tracking — and the thing customers notice when it is missing.