Engineering & Infrastructure

Retry Logic and Idempotency for Remittance APIs: A Developer's Guide

Idempotency key design, unknown outcomes, retry classification, backoff with jitter, webhook ordering and reconciliation for money movement

A timeout on a payment or payout call does not tell you the request failed. It tells you that you stopped waiting. The provider may have acted, and the code that runs next has to find out before it does anything that moves money a second time.

This is the developer companion to retrying payments and payouts safely, which covers the operations side: who decides, what customers are told, how exceptions are handled. Here the focus is code and contracts: idempotency key design, calls with unknown outcomes, which errors to retry, how to space retries, webhooks that arrive twice or out of order, and why reconciliation is still needed when all of that works.

01 · THE CONTRACT

Retries and idempotency solve different problems

Retry logic decides whether and when to send a request again. Idempotency decides what the receiving system does when the same logical request arrives more than once. A retry policy without idempotency generates duplicates; idempotency without a retry policy leaves transfers stuck after every network blip.

HTTP gives you less than you might hope. RFC 9110 defines PUT and DELETE as idempotent and POST as not, but "create a payout" is almost always a POST, and the guarantee you need is about business effect. It has to be written into the API contract with every provider you call and every client that calls you.

Your idempotency controls also stop at your own database. Inside a partner's system you have whatever the partner offers: a documented idempotency key that replays the original result, a unique client reference it rejects if seen before, or nothing. Find out which, per connection. With nothing, every resend after an unknown outcome is a potential duplicate, and a status lookup or a human check must come first.

02 · KEY DESIGN

Designing idempotency keys properly

If you expose an API that creates transfers or payouts, this is what to build. If you consume one, this is what to ask the provider before writing a retry loop.

One key per logical operation

The client generates the key once, when it decides to act, and stores it with the action before the first attempt. Every retry sends the same key; a genuinely new action, such as paying with a different card, gets a new one. The common bug is generating the key inside the HTTP client, so each retry carries a fresh key and the server sees three separate requests.

Use a random value with enough entropy that collisions are not a concern (a version 4 UUID is typical), or derive it from your own identifiers, such as payout ID plus attempt number. A derived key survives a crash.

Scope the key

Scope keys at least to the API credential, ideally to the endpoint too, so two clients cannot collide and a payment key cannot be confused with a payout key. Enforce uniqueness with a database constraint on (scope, key), not a check-then-insert that two concurrent requests will both pass.

Store more than the key

  • A request fingerprint: a hash of method, path and normalised body (sorted keys, consistent number formats), so harmless reformatting does not look like a different request.
  • A state: in progress, completed, or failed before any side effect.
  • The original response: status code and body.
  • The resource created, such as the payout ID, plus created and expiry times.

Decide what a repeat returns

How a server should answer a key it has seen
SituationResponseWhy
Same key, same fingerprint, completedReplay the stored status and body exactlyThe client gets the answer it missed; nothing runs twice
Same key, same fingerprint, in progress409, retry laterStops a second worker running the operation in parallel
Same key, different fingerprint409 or 422, never processedA client bug; executing it would attach a different payout to the old key
Same key, failed before any side effectAllow the retryA validation error should not lock the key for good

Replay the original body, not a "duplicate" error or a fresh timestamp. Client code that branches on the response would otherwise behave differently on retry, the path least likely to have been tested.

Keep keys at least as long as any client might retry, plus a margin. Twenty-four hours is a common floor for payment APIs; payouts that can sit with a partner over a weekend justify longer. A key that expires mid-schedule removes the protection exactly when it is needed.

Write the record and the effect together. Insert the idempotency record and the payout row in one database transaction. If the process dies between "payout created" and "key marked complete", a retry finds an in-progress key and no way to know the payout exists.

03 · UNKNOWN OUTCOMES

The timeout problem: look up before you resend

Every outbound call ends in a definite success, a definite failure, or an unknown. Code that collapses these into two will eventually pay someone twice. Give the unknown its own state on the record, such as submission_unknown.

  • Failure before the request left (DNS failure, connection refused, TLS failure): the provider cannot have acted. Safe to retry.
  • Failure after it was sent (read timeout, reset mid-response, a gateway 504): the provider may have acted. Unknown.
  • A 2xx you could not parse: the provider almost certainly acted, but you have lost its reference. Unknown.

Many HTTP libraries raise the same exception for "could not connect" and "connected but no reply"; wrap them so callers get a distinct result for each. Then:

  1. Persist the unknown state first, so no scheduler, user action or second worker creates a replacement.
  2. Look up status by your own reference. Every outbound request should carry a client reference the provider stores and can search on.
  3. Wait, and look more than once. Some providers commit asynchronously, so an immediate lookup can return "not found" for a request about to exist.
  4. Act on the answer. Found: adopt the provider's state and ID; do not resend. Confirmed absent: resend with the same reference and key. Unclear, or no lookup endpoint: hand the item to a person.

"Not found" is only trustworthy if the provider guarantees accepted requests become visible to lookups within a known time. If the documentation is silent, ask.

04 · CLASSIFICATION

Classifying failures: what to retry and how

A retry policy maps failure types to actions. Write it per provider, because status codes are used inconsistently. This is a sensible default for money-moving calls.

Retry classification for payment and payout API calls
ResultRetry?Action
Connection refused, DNS or TLS failureYesBackoff with jitter, same key
Read timeout or reset after sendingNot blindlyMark unknown; look up by reference; resend with same key only if confirmed absent
400 or 422 validation errorNoFix the data; the corrected request is a new operation with a new key
401 or 403NoCredentials or permissions; alert, do not loop
409, key reused with a different bodyNoClient bug; investigate before anything else is sent
Duplicate reference rejectedNoThe first request landed; look it up and adopt its state
429 with Retry-AfterYesWait at least Retry-After, same key; slow the whole queue for that provider
500With careUnknown unless the provider documents that a 500 means nothing was saved; look up first
502 or 503YesUsually never reached the application; backoff, same key, retry limit
504Not blindlyThe gateway gave up; the application behind it may have finished; treat as unknown

Business refusals, such as insufficient funds, an invalid account or a compliance rejection, are never retried automatically; they fail identically the second time. And when many calls to one provider fail together, stop sending for a while rather than letting every transfer retry on its own schedule. A circuit breaker that opens after a run of 5xx responses does this.

05 · BACKOFF

Exponential backoff with full jitter

Exponential backoff spaces attempts further apart each time. Jitter randomises each delay so that a thousand clients that failed together do not retry together. Full jitter picks the delay uniformly between zero and the exponential ceiling, which spreads load better than adding a little noise to a fixed delay.

BASE = 0.5 seconds
CAP = 30 seconds
MAX_ATTEMPTS = 6

function submit_with_retry(operation):
    key = operation.idempotency_key        # created once, stored with the operation
    for attempt in 0 .. MAX_ATTEMPTS - 1:
        result = send(operation.request, key)

        if result.succeeded:
            return result
        if result.outcome_unknown:
            status = lookup_by_reference(operation.reference)
            if status.found:
                return adopt(status)       # it landed; never resend
            if not status.confirmed_absent:
                return hand_to_operator(operation)
        else if not retryable(result):
            return result                  # business refusal or our own bug

        delay = random_uniform(0, min(CAP, BASE * 2 ^ attempt))
        if result.retry_after:
            delay = max(delay, result.retry_after)
        sleep(delay)

    return hand_to_operator(operation)     # retry limit reached
  • Retry from a queue, not a sleeping request handler. Schedule the next attempt as a delayed job carrying the same key, so a restart loses nothing.
  • Claim work atomically. Two workers can pick up the same payout; claim with a conditional update so exactly one wins.
  • Cap total time as well as attempts. A quote or payment session may expire before retries run out.
  • Budget retries globally. Retries piled on an outage prolong it.
06 · WEBHOOKS

Webhooks: store first, deduplicate, guard the state machine

Providers deliver webhooks at least once, not exactly once. Expect duplicates, delays and events out of order.

Store before processing

Verify the signature, write the raw body and headers to durable storage, return 2xx, then process asynchronously. If processing throws, the message still exists and can be replayed after the fix. A handler that applies the event inline and returns 500 on failure depends on the provider's retry schedule, which may give up before you deploy.

Deduplicate by event ID

Put a unique constraint on the provider's event ID, scoped to the provider, so a second delivery fails the insert and is acknowledged without being applied. Where there is no event ID, build one from the provider reference, the new status and the event timestamp.

Guard against out-of-order events

A "processing" event can arrive after "paid". A naive handler moves the payout backwards, and the next customer notification is wrong. Define the payout state machine explicitly and apply an event only if the transition from the current state is allowed; paid, returned and cancelled are terminal except through a deliberate reversal. If the provider sends sequence numbers or event timestamps, ignore anything older than the last event applied.

Record rejected events with the reason. "Paid, then returned three days later" is real and must be applied; "paid, then processing" is noise. For terminal events that move money in your books, consider fetching the resource from the provider's API before applying it, so a stale or malformed payload cannot complete a payout on its own.

07 · RECONCILIATION

Reconciliation is the final safety net

Keys, lookups and state guards reduce mismatches; none removes them. A webhook may never be sent, a lookup may lag, a provider may fix something by hand without telling you. Reconciliation compares your records with the provider's report, independently of the request path, at least daily. Match on your reference first, the provider's ID second, and sort every difference: in their records but not yours (often an unknown outcome resolved the wrong way), in yours but not theirs, or in both with a different amount or status.

Each difference becomes a case for a person. Automatic correction from a reconciliation job is how one error becomes two. See payment reconciliation for cross-border money transfers.

08 · SCENARIO

A worked scenario: a payout submit times out

Timings are illustrative. A transfer's payment has cleared and its payout is ready for a recipient's bank account.

  1. 10:00:00, claim. A worker claims the payout with a conditional update. Its key and client reference, derived from the payout ID and attempt number, were stored at approval.
  2. 10:00:01, submit. The create-payout request goes out with the client reference; the outbound log records it with a correlation ID.
  3. 10:00:31, timeout. The body was sent but no reply came in 30 seconds. The client returns "outcome unknown", and the payout moves to submission_unknown.
  4. 10:00:41, first lookup. By client reference: 404.
  5. 10:01:11, second lookup. The payout is there, status "processing", with the provider's ID. The original request was accepted; the reply was lost.
  6. 10:01:11, adopt. The worker stores the provider's ID, moves the payout to sent and schedules status checks. Nothing is resent.
  7. 10:14:02, webhook. "Paid" arrives, is stored, deduplicated and applied: sent to paid is allowed.
  8. 10:14:05, late event. A delayed "processing" event is stored, but the guard refuses paid to processing and logs why.
  9. Next morning, reconcile. The provider's report shows one payout for that reference, matching.

What made it safe: "no reply" was not treated as "failed", the first "not found" was not trusted alone, the reference was searchable, and the state machine ignored the late event. A resend at 10:00:41 to a provider that does not deduplicate references would have paid the recipient twice.

09 · REMITSO

Doing it with RemitSo

RemitSo's admin panel handles the platform side of these patterns; your team still owns the decisions, the partner relationships and anything you build on the API. See the admin features.

  • Outbound API log: every call to a partner logged with request and response, status, error, duration and correlation ID, keys masked, filterable to failed calls: the evidence for whether a timed-out request left the platform.
  • Webhook logs: every incoming message saved before processing, with type, provider, status, attempts, error and time received; raw message and failure reason one click away; retry in one click. Reading a provider's message and re-applying it are separate permissions.
  • Failed background jobs, listed under system utilities for review and retry.
  • Payout errors sorted by fault: a fault on the platform's side or a missing setting leaves the payout in Error for your team to resend, and the customer is not told it failed; customers hear a payout failed only when the partner refused it. A payout the partner may already hold is never resent without an operator confirming.
  • Payouts claimed atomically by one process, and asked about on a schedule; one the partner stops answering about appears on Scout and is never cancelled. No payout is sent before the transfer's payment has cleared.
  • Duplicate identity results recorded as Duplicate and not retried.
  • Plain-language alerts to subscribers: "Payout not sent because of a fault on our side" says whose side the fault was on and whether the partner must be asked before retrying; "Duplicate payment captured" and "Payment received on closed attempt" surface payment duplicates.
  • Rate limiting on sign-in, sign-up, verification codes and password resets in the apps and console.
  • API documentation: the Client API used by the mobile apps and web client is documented live inside the console and kept in step with the API.

Before building against any remittance API, see evaluating a remittance API before you sign; for the operations view of failed partner calls, see reading API and webhook logs without an engineer. Or book a demo to walk through a timed-out payout in the console.

FAQ

Frequently asked questions

Should the client or the server generate the idempotency key?

The client, because only it knows when two requests are attempts at the same action. The server enforces uniqueness within a scope and stores the result against the key.

Is a 500 response safe to retry?

Not by default. A 500 can be raised after the operation was saved. Unless the provider documents otherwise, treat it as an unknown outcome and look up status first.

Why return 2xx to a webhook before processing it?

Because the message is already stored. The provider stops redelivering, and your processing and replays work from the stored copy rather than the provider's schedule.

If we have idempotency keys, do we still need reconciliation?

Yes. Keys do not catch a missed webhook, a lagging lookup or a manual change on the provider's side. Reconciliation compares the records directly.

Built by people who have helped MSBs for years.

The risk checks on every online transfer come from what they see every day.

  • The person paying isn't the customer
  • One bank account, several customers
  • A disposable email address
  • Sign-in from a high-risk location
  • The same person signing up twice
  • A name close to a sanctions list
Book a demo

See RemitSo running with your corridors.

  • 30 minutes with our team, on video
  • The real admin panel and customer apps
  • White label or source code, explained with pricing
  • Your compliance and launch questions answered
Loading the form…

Video