Engineering & Infrastructure

When a Payout Partner Fails: Reading API and Webhook Logs Without an Engineer

A practical guide for operations teams on diagnosing failed payouts, deciding whose problem it is and retrying safely

When a payout fails, the worst position for an operations team is guessing. The customer is waiting, the recipient is asking, and the only person who can explain what happened is an engineer who is busy with something else. It does not have to work that way. Most partner failures can be read, classified and resolved by the operations team itself, provided every call to the partner and every message back from it is recorded in a form people can actually read.

This guide is for operations leads and the people who handle payout exceptions day to day.

RemitSo is one way to run this. Its admin panel logs every outbound partner request and response, with payloads, headers, status code and the raw error, keys masked, and saves every incoming webhook before processing. Alongside the logs, alerts tell your team in plain language what went wrong and whose side it is on. It does not decide for you: your team still confirms the diagnosis, talks to the partner and owns the retry. This guide aims to make that judgement quick and repeatable on any platform.

01 · THE LOG ENTRY

How to read an outbound API log entry

An outbound API log entry is the record of one request your platform made to a partner and what came back. For payouts, that is usually a request to create a payout, check its status or cancel it. A useful entry answers five questions without anyone opening a code editor.

The error message

The partner's explanation of what went wrong, exactly as returned: "beneficiary account number invalid", "insufficient prefunded balance", "authentication failed". Logs record this raw technical text, not a translation, and that is what you want: it is the exact wording a partner's support team will search for. Read it first; it often answers the question on its own.

The status

The standard signal of what kind of outcome the call had. Operations teams only need the broad families:

  • 2xx: the partner accepted the request. If the payout still looks wrong, the problem is later in the journey, not at submission.
  • 4xx: the partner rejected what was sent. The request itself was the problem: bad credentials, missing or invalid fields, a duplicate reference, a payout the partner will not accept.
  • 5xx: the partner had a problem processing a request that may have been perfectly valid. Outages, overloaded services and internal failures sit here.
  • No status (timeout): no answer came back in time. This is the most uncertain case, because the partner may or may not have acted on the request.

The duration

How long the call took. Fast failures usually mean the partner looked at the request and refused it. Slow failures and timeouts usually point to the partner's systems being under strain or unavailable.

The correlation ID

A unique reference that lets both sides find the same request in their own records. It is the most useful thing you can give a partner's support team: without it they search by time and amount; with it they go straight to the request.

Masked keys

Enough of the credential to confirm which one was used, never the whole secret. Operations staff can read logs without seeing live credentials, and entries can be shared with a partner without leaking access. If the masked key on a failing call is not the one you expect, a key change may not have been applied everywhere.

Rule of thumb: read the error message first, the status second and the duration third. Between them, they usually tell you whose problem it is before you have opened anything else.

02 · WEBHOOKS

Why incoming webhooks must be saved before they are processed

Outbound calls are half the conversation. The other half is the partner telling you, usually by webhook, that a payout was paid, rejected or returned by the receiving bank.

The risk with webhooks is quiet loss. If a platform processes a message the moment it arrives and something goes wrong, such as an unexpected field format, a poorly built system may simply drop it. The partner thinks it told you the payout was paid; your platform still shows it as pending; nobody knows until the customer calls.

Saving every message before processing closes that gap. Even if processing fails, the message exists, and someone can read why it failed and process it again once the cause is fixed.

In RemitSo, the webhook log in the admin panel works this way. Every incoming message is stored first; the raw message and the reason it failed are one click away; and it can be retried in one click. For an operations team, the benefit is certainty: a partner update cannot vanish because the platform choked on it.

Watch for: a partner insisting it sent a status update that your platform never reflected. Before escalating, check the webhook log. If the message is there with a processing error, the fix is on your side and the partner did its job.

03 · OURS OR THEIRS

How to decide whether a failure is ours or the partner's

The most valuable decision on a failed payout is also the simplest: who needs to act? If the cause is on your side, your team fixes it. If it is the partner's, the job is to tell them precisely what happened, while keeping the customer informed. Use the table below as a starting point and extend it with the messages your own partners return.

Reading a failed partner call: likely owner and next step
What the log showsLikely ownerTypical causeNext step
4xx, "authentication failed" or "invalid credentials"OursExpired or rotated key, wrong environment configuredCheck the masked key against the current one; update configuration; retry
4xx, "invalid account" or "beneficiary name mismatch"Ours (customer data)Recipient details entered incorrectlyContact the customer for corrected details; do not resend the same instruction
4xx, "missing field" or "invalid format"OursRecipient form or mapping does not match what the partner requiresFix the data or configuration; involve an engineer only if the mapping itself is wrong
4xx, "insufficient balance" or "prefunding exhausted"Ours (treasury)Prefunded account at the partner has run lowTop up the partner balance; retry once funds are confirmed
5xx, fast failure, many payouts at onceTheirsPartner service faultNotify partner with correlation IDs; hold retries until they confirm recovery
Timeout, long durationUnknown until checkedPartner slow or unavailable; request may or may not have been receivedCheck status with the partner before any retry
2xx on submission, later rejection webhookUsually theirs or the receiving bank'sDownstream rejection after acceptanceRead the webhook reason; treat as a returned payout
Webhook received but failed processingOursPlatform could not apply the updateRead the failure reason; fix; retry the webhook

Two patterns hold. A single failure is usually about that transaction; many with the same message in a short window are usually about the connection or the partner. And a timeout is neither ours nor theirs until you know whether the partner received the request.

Use the log and the alert together

The log and the alert answer different questions. The log is the evidence: the exact request, the response, the status code and the error text, which is what you quote to a partner or an engineer. The alert is the diagnosis: a plain-language statement your team can act on without decoding a status code. In RemitSo, for example, the alert "Payout not sent because of a fault on our side" names whose side the fault was on and whether the partner must be asked before the payout is retried. Read the alert to decide what to do; open the log to prove it and to brief whoever has to fix it.

04 · ESCALATION

What to send a partner when the failure is theirs

A vague report ("payouts are failing, please check") gets a request for more information. A precise one gets investigated.

A good escalation contains:

  1. Correlation IDs: a representative handful, plus the total count.
  2. Time window, with time zone, and whether it is still happening.
  3. Status and error message exactly as returned.
  4. Duration pattern: "timing out after 30 seconds" beats "failing".
  5. Scope: which corridor, payout method or endpoint is affected, and which work.
  6. What you have ruled out, such as credentials being current.

Because RemitSo masks keys in the outbound log and records the full request and response, the exact payload and error a partner needs can be shared without redacting screenshots first.

Template: "Between 14:05 and 14:40 UTC, 18 payout creation calls returned 503 'service unavailable' in under 200 ms. Sample correlation IDs below. Status checks on the same credentials are succeeding. Please confirm whether these requests were received and when service is restored."

05 · SAFE RETRY

How to retry without creating a second problem

The two common mistakes are retrying something that will fail the same way again, and retrying something that may already have succeeded.

  • Retry only when the cause is resolved. Resending an instruction with an invalid account number just adds another rejection. Fix the cause first: corrected details, renewed credentials, a topped-up balance, or the partner's confirmation that service is back.
  • Never retry a timeout blind. A timed-out request may have been acted on. Check status with the partner, or wait for a webhook, first. This is how duplicate payouts happen.
  • Never edit an approved payout. What goes out should be exactly what was approved. If details must change, ask the partner to stop or return the payout where they can, and once it is cancelled or returned, create a new one that goes through approval again. In RemitSo, payouts are locked once approved and are never edited, so nobody can quietly change where money goes after sign-off.
  • Retry from the log, not by workaround. A one-click retry keeps the original record, the failure and the retry linked. Recreating a transaction by hand breaks that chain and makes reconciliation harder.
06 · SCENARIO

Scenario: a Tuesday afternoon with two kinds of failure

The numbers below are illustrative, chosen to show the reasoning rather than to describe any real operator.

An operator pays out to bank accounts and mobile wallets through one partner. At 14:10, the operations lead sees items flagged as failing partner calls, and a run of payout alerts saying the payouts were not sent because of a fault on the partner's side.

Step 1: look at the pattern, not the first item

The outbound log shows 26 failed payout calls since 14:05, in two groups:

  • 23 calls: status 503, "service unavailable", each failing in under 200 milliseconds, all on bank account payouts.
  • 3 calls: status 400, "beneficiary account number invalid", spread across the afternoon, on unrelated transfers.

Step 2: classify each group

The 23 fast 503s, all on one payout method and starting together, point to the partner, which matches what the alerts said. The log confirms it: wallet payouts on the same credentials are succeeding, which rules out authentication on the operator's side. The 3 400s are customer data problems; the raw error in each entry names the field.

Step 3: act on what is yours, report what is theirs

For the 3 invalid-account payouts, which the partner refused, the team contacts the customers for corrected details; each refused payout is cancelled rather than edited, and a new payout with the corrected details goes through approval normally. In parallel, the lead sends the partner the time window, count, status and message, five correlation IDs, and the note that wallet payouts are unaffected. No engineer is involved, because nothing on the operator's side is broken.

Step 4: retry once the partner confirms

At 14:52 the partner confirms service is restored and that none of the 23 requests were processed. The team retries them. Twenty-two go through; one returns "beneficiary name mismatch" and joins the customer-data queue.

Step 5: check the webhooks

Paid confirmations arrive over the next hour. One fails processing; the webhook log shows the raw message and a field format the platform did not expect. Now an engineer is involved, with a precise question and the exact message. After the fix, the webhook is retried in one click and the payout shows as paid.

07 · ESCALATING INTERNALLY

What to check before calling an engineer

Engineers should be involved when they are genuinely needed, not as the first step on every red row. Before escalating internally, operations should know whether it is one transaction or many, whether the log shows a 4xx, a 5xx or a timeout, whether the partner has been told, and what any webhook failure reason says.

Genuine reasons to call an engineer include a partner changing its message format, a field mapping that is wrong for a whole type of transaction, and a webhook that keeps failing for a reason configuration cannot fix. On RemitSo's white label plans, engineering time is billed only for time logged, another reason to arrive with the evidence gathered; the pricing page sets out the rates.

08 · CHECKLIST

A checklist for partner failures

  1. Look at the pattern: how many failures, since when, which payout method and corridor.
  2. Read the alert for the plain-language diagnosis, then the log for the error message, status and duration.
  3. Classify: ours, theirs, or unknown (timeouts).
  4. For ours: fix the cause (data, credentials, balance, configuration) before any retry.
  5. For theirs: send correlation IDs, time window, status, message and scope.
  6. For unknowns: confirm with the partner whether the request was received before retrying.
  7. Never edit an approved payout. If details must change, have the partner stop or return it, then send a correct payout through approval.
  8. Retry from the log so the history stays linked.
  9. Check the webhook log for updates that arrived but failed processing.
  10. Bring in an engineer with a specific question and the evidence attached.
09 · REMITSO

Doing it with RemitSo

RemitSo's admin panel gives operations teams the records described above, so most partner failures can be handled without waiting for an engineer. Your team still makes the decisions and owns the partner relationship.

  • Outbound API log: every call to a payout, payment or identity provider is on the record, each request and response with payloads, headers, status code and the raw error, filterable to failed calls, so the team has the exact evidence to quote. See the admin features.
  • Masked keys: entries can be read by operations and shared with partners safely.
  • Webhook logs: every message saved before processing, with type, provider, status, attempts, error and time received, so no partner update is lost; raw message and failure reason one click away; retry in one click.
  • Plain-language alerts: alerts such as "Payout not sent because of a fault on our side" say whose side the fault was on and whether to ask the partner before retrying, and go to the people subscribed to them, so nobody watches a screen all day.
  • Scout: measures every five minutes and groups payouts needing action, and in system health shows partners failing calls and provider messages never applied, so problems surface before customers report them.
  • Every payout chased on a schedule: each payout is followed up on a schedule that fits it, so one the partner never answers about reaches an operator instead of sitting silently.
  • Holds recorded on the payout: a hold agreed with the partner can be recorded on the payout, with a required reason that is kept, so the next person on shift sees why it is waiting.
  • Payouts locked once approved: never edited, which protects against mistakes and tampering.
  • Roles and sign-in records: each person sees only what their role allows, and every staff sign-in is logged with time, IP address and whether two-step sign-in was completed.

To see how these fit with the rest of the platform, read the release notes or book a demo.

FAQ

Frequently asked questions

What is a correlation ID and why does the partner need it?

It is a unique reference attached to a specific request. Both sides can use it to find the same call in their own records, which turns a partner investigation from a search by time and amount into a direct lookup.

Is it safe to retry a payout that timed out?

Not until you know whether the partner received it. A timed-out request may have been processed. Check status with the partner or wait for a webhook before sending it again, or you risk paying twice.

Why can't we just correct the recipient details on an approved payout?

Because the approval covered specific details. Changing them afterwards would mean money goes somewhere that was never approved. Having the partner stop or return the payout, then sending a correct one through approval, keeps the approval meaningful and the history clear.

What happens if a partner's webhook fails to process?

If the platform saves messages before processing, as RemitSo does, the message remains in the webhook log with the reason it failed. Once the cause is fixed, it can be retried and the transaction updated.

Built by people who have helped MSBs for years.

The risk checks on every online transfer come from what they see every day.

  • The person paying isn't the customer
  • One bank account, several customers
  • A disposable email address
  • Sign-in from a high-risk location
  • The same person signing up twice
  • A name close to a sanctions list
Book a demo

See RemitSo running with your corridors.

  • 30 minutes with our team, on video
  • The real admin panel and customer apps
  • White label or source code, explained with pricing
  • Your compliance and launch questions answered
Loading the form…

Video