Skip to content
All writing

TypeScript and API design

Designing Idempotent Third-Party Integrations That Recover Cleanly

A production-minded method for making provider calls, webhooks, retries, reconciliation, and operator recovery converge on one trustworthy result.

Mubashir Hussain10 min read

Reliable integrations are not created by finding the perfect HTTP client. They are created by accepting an uncomfortable fact: the local application and the provider will sometimes disagree about what happened. A request can complete after the client times out. A webhook can arrive twice, arrive late, or arrive before an earlier event. A database commit can succeed while a follow-up notification fails. Once money, identity, DNS, mailboxes, or deliveries are involved, treating these cases as unusual exceptions turns ordinary network behavior into customer-facing corruption.

The goal is therefore not exactly-once transport. Most systems cannot promise that across independent databases and networks. The practical goal is convergence: every repeated command, delayed event, scheduled repair, and operator action should move the system toward the same correct business outcome. Work on Be My Courier and Custom Name Domain makes this concrete because those products coordinate provider-owned facts with local workflows. The design has to preserve provider authority while still giving users fast, understandable progress.

Start with a business operation identity

An idempotency key is useful only when it represents a stable business operation. Generating a fresh random key for every retry defeats the purpose because the provider sees every attempt as new. Reusing one key for unrelated inputs is also unsafe because it hides a programming error behind an old response. The application needs an operation record created before the first outbound attempt, with a durable identifier and an immutable description of intent.

For a marketplace payout, the identity might be tied to the booking and payout transition. For domain registration, it might combine the order, domain, and registration attempt. The important part is that all retry paths can find the same record. A browser retry, a queue retry, an operator retry, and a recovery job should not invent competing identities.

Stripe documents that an idempotency key lets a client safely repeat a creation or update request and receive the stored result. It also compares parameters when a key is reused. That behavior suggests a local invariant: once an operation starts, material parameters are frozen. If the requested amount or destination changes, create a new business operation through an explicit transition rather than silently mutating the old one.

A useful operation record contains the local operation ID, provider, operation type, target resource, normalized request fingerprint, provider request ID when available, attempt count, last outcome, and timestamps. It is not merely a log line. It is the durable coordination point that every execution path consults.

Separate command acceptance from provider completion

Many integration bugs begin when one application status tries to represent several different facts. A local request can be accepted while the provider result remains unknown. It can be known at the provider while local projections are unfinished. It can be complete while a secondary email or notification is still pending. Collapsing all of that into success and failure encourages code to lie.

Model the stages explicitly. A command may be accepted, dispatching, awaiting provider confirmation, confirmed, applying local effects, completed, retryable, or blocked for review. These names are examples, not a universal enum. The useful test is whether each state tells the next worker or operator what is safe to do.

An HTTP timeout after sending a payment request is not proof of failure. It is an unknown outcome. The safe next step is to retry with the same operation identity or query the provider, not to create another payment. A provider validation error, by contrast, may be a terminal rejection that requires corrected user input. A rate limit is normally retryable. Authentication failure is usually operationally blocked. The state machine should preserve those distinctions.

Return a stable local operation resource to the UI as soon as intent is recorded. The page can then show pending truthfully and poll or receive updates. This keeps the request latency bounded without pretending a cross-provider workflow is synchronous. It also lets support staff reference the same operation the workers use.

Treat webhooks as notifications, not ordered commands

Webhook delivery is at-least-once in practice. Stripe explicitly warns that endpoints can receive duplicate events and that event ordering is not guaranteed. A robust handler therefore verifies authenticity, stores a receipt, acknowledges quickly, and schedules idempotent processing. It does not assume that receiving an event means the whole business transition should run inline.

Use the provider event ID as a unique receipt key where the provider guarantees uniqueness. Insert the receipt under a database unique constraint. If the insert conflicts, return a successful acknowledgement because the event is already known. This turns duplicates into normal no-op outcomes rather than errors that cause more retries.

Deduplicating delivery is necessary but not sufficient. Two distinct provider events can describe overlapping state, and an older event may arrive after a newer one. The processor should retrieve the current provider object when correctness depends on its latest state, then apply a monotonic local transition. Event payloads remain valuable evidence, but they are not always authoritative snapshots.

Keep acknowledgement and processing separate. Signature verification and durable receipt should be fast. Provider lookups, database projection updates, email, and downstream jobs can happen asynchronously. This reduces timeout-driven redelivery and isolates optional effects. If a worker fails after the receipt is stored, the receipt remains visible and retryable.

Make the provider authoritative about provider facts

Local models need clear ownership rules. The provider owns facts such as whether a charge succeeded, an external account is enabled, a domain registration exists, or a mailbox has been provisioned. The application owns product facts such as whether prerequisites were accepted, which user requested an action, and whether local business policy allows the next transition. Confusion arises when local code overwrites provider truth with an optimistic assumption.

For every integrated field, decide whether it is command intent, cached provider state, or local business state. Intent records what the application asked for. Cached provider state records the last observed external truth and when it was observed. Business state combines evidence under product rules. Keeping these concepts distinct supports honest UI and safer repair.

Suppose a courier delivery confirmation makes a payout eligible. Eligibility is local policy; payout completion is provider truth. The local system may transition the payout operation to awaiting confirmation after sending the command, but should not label it paid until the provider confirms it. Likewise, a domain checkout may be paid while registration remains pending. Showing those as separate stages is more accurate than a single order status.

Provider authority does not mean blindly accepting any callback. Verify signatures, account context, object ownership, currency, amount, and expected relationships. Authority answers who knows the external outcome, not whether an incoming message bypasses product invariants.

Retry with budgets, classification, and jitter

Retries are a reliability tool only when the operation is idempotent and the error is classified. Retrying every failure can amplify outages, exhaust provider quotas, and keep permanently invalid work alive. Define a policy for each operation rather than hiding one generic retry loop in an HTTP wrapper.

Transient network failures, rate limits, and selected server failures may be retryable. Validation errors and forbidden operations usually are not. An ambiguous connection failure after dispatch requires the same idempotency key. A conflict may require refreshing provider state before deciding. Stripe describes exponential backoff for retries and explains how idempotency supports reconciliation after uncertain outcomes; the local implementation should add bounded attempts and randomized jitter so workers do not synchronize during recovery.

Record the next attempt time and reason. A queue message alone is not durable enough if it can disappear or be duplicated. Workers should claim due operations with a lease or transactional update, perform one bounded attempt, and persist the result. If a worker crashes, the lease expires and another worker continues from the same operation record.

Set separate limits for request duration, total attempt count, and total recovery age. After the automatic budget is exhausted, move the operation to an operator-visible blocked state instead of retrying forever. The operation still retains enough evidence for a deliberate resume.

Reconcile because webhooks are not a complete ledger

Even a well-built webhook endpoint can miss events during configuration mistakes, deployment outages, provider incidents, or secret rotation. Reconciliation is the independent path that proves local projections still agree with the provider. It should exist from the beginning for high-value integrations, not be invented during an incident.

There are two useful forms. Targeted reconciliation examines operations in unknown, pending, or stale states and retrieves their provider objects. Broad reconciliation scans a bounded provider time window or resource set and compares it with local records. Targeted checks are cheaper and frequent; broad checks catch missing local operations and systemic drift.

The reconciliation algorithm should be idempotent. Fetch external truth, validate ownership, compare normalized state, apply only allowed transitions, and record what changed. Do not replay arbitrary historical side effects. Notifications, rewards, or downstream actions should have their own unique effect keys so a repaired transition cannot issue them twice.

Store checkpoints for paginated scans and overlap time windows to avoid gaps. Overlap produces duplicates, which is safe when receipts and effects are idempotent. A gap is harder to detect. Expose reconciliation lag, mismatch counts, repaired counts, and permanently unresolved items as operational signals.

Design every side effect as its own recoverable step

A provider confirmation often triggers several local effects: update an order, create a transaction, send an email, notify a socket room, award loyalty points, or unlock the next workflow. Wrapping a database update and a network call in one function does not make them atomic. If the process crashes between steps, a retry may duplicate one effect or skip another.

Commit authoritative local state and durable outbox entries in one database transaction. An outbox worker then delivers each external or asynchronous effect. Give every effect a unique key based on the business transition and effect type. The email for payout completion, for example, should have one identity independent of how many times the transition processor runs.

Keep optional effects from blocking core truth. A notification provider outage should not revert a confirmed payment. The UI can derive completion from the authoritative transaction while the notification remains retryable. Conversely, do not mark a provider operation complete merely because an email succeeded.

This structure also improves testing. A transition test can assert the database state and expected outbox records without depending on live providers. Worker tests can prove duplicate delivery is harmless. Recovery tests can stop execution after each boundary and confirm that a retry converges.

Give operators evidence and constrained recovery actions

Some integration failures need human judgment: provider account suspension, mismatched ownership, unexpected currency, irreconcilable legacy data, or a policy exception. The admin surface should explain the operation without exposing secrets. Show the business resource, current state, last provider observation, attempt history, redacted request identifiers, error classification, and next safe action.

Recovery controls should invoke the same domain commands as automation. A Retry button should renew the existing operation, not bypass idempotency. A Refresh provider state action should reconcile, not overwrite. A Resolve action should require a reason and produce an audit entry. Dangerous actions such as creating a replacement payment need explicit confirmation and a new operation identity.

Observability should connect logs, metrics, and traces through the local operation ID and provider request or event IDs. Measure pending age, ambiguous outcomes, webhook receipt latency, processing latency, duplicate rates, retry counts, reconciliation drift, and blocked operations. Alerts should point to a recoverable queue, not merely announce that an endpoint returned errors.

Define readiness as recoverability

An integration is not production-ready because the happy path passed once. It is ready when duplicate commands are harmless, uncertain outcomes remain truthful, delayed events converge, provider truth can repair local drift, optional side effects cannot corrupt core state, and an operator can understand what happened.

Test those properties deliberately. Inject a timeout after the provider accepts a request. Deliver the same webhook twice. Deliver events out of order. Crash after the database commit but before queue acknowledgement. Disable a notification provider. Run two workers against the same operation. Expire a lease. Reconcile a missing webhook. For each case, assert one business outcome and complete evidence.

The broader design principle is simple: make retries boring and ambiguity visible. A durable operation identity, explicit states, provider-authoritative reconciliation, isolated effects, and constrained operator tools turn unreliable networks into manageable product behavior. That is what allows integration-heavy systems to grow without making every vendor incident a data-repair project.

Primary sources

  1. 1.Idempotent requests — Stripe
  2. 2.Receive Stripe events in your webhook endpoint — Stripe
  3. 3.Advanced error handling — Stripe

Portfolio evidence

Related writing