Payments are the moment a digital promise becomes tangible. That is why reliability matters more than cleverness.
Over three integrations we learned the same lesson three times: a payment system is only as good as its failure handling. Here is what that actually looks like.
Lesson 1 — Make the state machine explicit.
Every payment has a life cycle: initiated, submitted, confirmed, completed, failed, and reversed. We modelled these as an enum with strict transitions. A payment in "confirmed" cannot jump to "failed" — it must pass through "reversed". This sounds obvious until you realize most bugs come from implicit states.
Lesson 2 — Retry like a human would.
Network timeouts are rarely permanent. We added exponential backoff with jitter, capped at three attempts, and a circuit breaker that pauses retries if the failure rate crosses 50% in a rolling minute. The key rule: never retry a customer-initiated action more than the user can see.
Lesson 3 — Tell the user what happened.
When a transaction fails, the message must name the cause. "Payment failed" is useless. "M-Pesa timed out — please check your phone and confirm, or try again" is actionable. We instrumented every failure with a code and surfaced it in the UI.
Our approach pairs clear transaction states with safe retries, useful observability, and copy that helps people understand what happened.
These patterns travel well beyond one provider. They are the foundation of every dependable commerce experience.
Build what comes next
Want to explore what these ideas could mean for your business? Start a conversation with our team.