Skip to content
All writing

Testing and production delivery

Concurrency-Safe State Transitions for Marketplaces and Real-Time Products

A practical guide to optimistic concurrency, transactions, constraints, leases, locks, retries, and idempotent completion in stateful products.

Mubashir Hussain10 min read

Concurrency defects rarely appear in a single happy-path request. They appear when two valid requests observe the same old state and both decide they are allowed to advance it. A buyer accepts while a provider cancels. Two workers finalize the same game. A webhook confirms payment while an operator issues a refund. A timer expires as a live event arrives. Each handler can look correct alone while their interleaving produces an impossible result.

Be My Courier, JSleeve, and Mining Access expose different versions of this problem. Marketplaces coordinate offers, payments, delivery, approval, and payouts. Real-time products coordinate socket events, timers, scheduled jobs, and persisted results. The solution is not to put a lock around the whole application. It is to define invariants, choose the smallest coordination primitive that protects them, and make every repeated completion harmless.

Concurrency safety is ultimately a data-model property supported by application code. The database must reject invalid races or make them detectable. Workers and providers then converge through retries using stable operation identities.

Define the transition as a compare-and-set operation

Represent important lifecycle changes as transitions from an allowed source state to one target state. A command is not update booking with these fields. It is accept offer if the booking is still proposed and the actor is the assigned buyer. The predicate is part of the write.

An atomic conditional update is often enough:

update bookings
set status = accepted, version = version + 1
where id = bookingId
  and tenant_id = tenantId
  and status = proposed
  and version = expectedVersion

If the affected count is zero, reload and classify the result. The operation may already be complete, the version may be stale, the actor may have the wrong scope, or another transition may have won. Do not report a generic server error when the database has revealed a meaningful conflict.

The version field makes optimistic concurrency explicit, but source-state predicates can also protect simple transitions. Use both when clients edit other mutable fields that must not be overwritten. Return the new version so subsequent commands state which snapshot they acted on.

Keep transition logic in one domain service shared by HTTP, sockets, webhooks, timers, and jobs. Separate handlers that each perform direct updates will eventually disagree about guards or side effects.

Put invariants in constraints, not preliminary queries

The pattern check then insert is unsafe under concurrency unless the database makes the check and insert atomic. Two requests can both observe no active assignment and then create one. A unique constraint on the real business key lets one win and the other receive a conflict that can be handled idempotently.

Useful constraints include one completion per game, one payout operation per booking, one active lease per resource, one processed receipt per provider event, and one transition effect per operation and effect type. Partial unique indexes can protect invariants that apply only to active states. Foreign keys and check constraints can prevent invalid relationships and values.

Design the key around business identity, not transport identity. Two webhook event IDs may represent the same provider object transition, while one event can be delivered many times. A receipt key deduplicates delivery; an effect key deduplicates the business consequence. Both may be required.

When a unique conflict occurs, load the winning record and decide whether it represents the same intent. If yes, return it as an idempotent success. If its normalized parameters differ, report a conflict. Do not swallow every constraint error because some reveal a genuine invariant violation that needs investigation.

Constraints are executable documentation. They protect every writer, including maintenance scripts and future services that bypass the original code path.

Choose transaction boundaries around one invariant

Use a transaction when several writes must become visible together for an invariant to hold. Accepting a marketplace offer may update the booking, reject competing offers, and create an outbox event. Completing a game may write the result, update standings, and create one notification intent. If partial visibility would be invalid, these writes belong in one transaction.

Keep transactions short. Do not call payment providers, send email, wait for sockets, or perform expensive computation while holding database locks. Record the state transition and an outbox item, commit, then deliver external effects. This reduces contention and prevents a provider timeout from extending the critical section.

MongoDB documents that multi-document transactions can make changes atomic but have operational and performance considerations; effective schema design can reduce how often they are needed. Embedding data that changes together can make a single-document write sufficient. In PostgreSQL, relational constraints and row locks support a different modeling style. Choose according to the invariant, not fashion.

Define retry behavior for transaction conflicts and transient aborts. The transaction function may execute more than once, so its in-transaction computation must not perform non-idempotent external work. Generate stable operation IDs before the transaction rather than inventing them on every callback attempt.

Use pessimistic locks only when waiting is the right behavior

Optimistic concurrency is effective when conflicts are uncommon and callers can reload. Pessimistic row locks are useful when a short critical section must serialize decisions on one resource. PostgreSQL documents that SELECT FOR UPDATE prevents competing writers and lockers from changing the selected row until the transaction ends.

Lock rows in a consistent order when a transition touches several resources. Inconsistent ordering can deadlock: one transaction holds booking A and waits for account B while another holds account B and waits for booking A. Databases can detect deadlocks and abort one transaction, but the application still needs bounded retry and observability.

Avoid locking broad tables for ordinary product transitions. Table-level locks reduce concurrency and can make a maintenance operation look like an outage. Ensure predicates are indexed so the database finds the intended rows quickly. Monitor lock wait time, transaction age, and deadlocks rather than assuming locks are free.

A lock does not replace authorization or state validation. After acquiring it, re-read the current row and apply the transition guard. Data may have changed while the request waited. A timeout acquiring the lock should return a retryable conflict or enqueue work rather than leaving an unbounded request.

Use database locks when the protected resource is in that database. Reaching first for a distributed lock around a database row often adds failure modes without strengthening the invariant.

Treat distributed locks as expiring leases

Background jobs and real-time coordinators sometimes need ownership beyond one database transaction. A lease records that worker A may act on resource R until a specific time. Acquisition must be atomic, renewal must prove ownership, and completion must verify that the lease is still current.

Include a unique random token or monotonic fencing token. If a worker pauses longer than the lease, another worker may acquire it. The old worker must not later delete or renew the new worker's lease. Redis documents safe release by comparing the stored unique value before deletion and notes the timing and availability assumptions behind distributed locking.

Fencing is stronger for downstream correctness. Each successful acquisition receives a higher token, and the protected storage rejects writes carrying an older token. This prevents a delayed former owner from committing after its lease expires. If the downstream system cannot enforce fencing, understand that a time-based lock alone cannot guarantee exclusivity during long pauses or clock and network failures.

Set lease duration longer than normal work but bounded enough for recovery. Renew only while making progress, limit renewal attempts, and surface stuck leases. On startup, scan expired running work and reconcile it. Never assume process liveness proves lease ownership.

Partition lease keys to the smallest resource needing exclusion. A global game-finalizer lock sacrifices availability and throughput when games are independent.

Make retries reuse the same business intent

Retries occur at many layers: users double-click, clients reconnect, queues redeliver, databases abort transactions, webhooks repeat, and operators press recovery buttons. Every path must either reuse the same business operation identity or prove the command is naturally idempotent.

Persist an operation record with its normalized parameters before external work. A marketplace payment intent may be keyed by booking and payment attempt. Game completion may be keyed by game and completion generation. The first execution creates the record under a unique constraint; later executions load it.

Stripe documents idempotency keys for safely retrying creation and update requests. Keep the provider key attached to the local operation. If the first call times out, reuse it. A new key means a new provider operation and should require an explicit business decision, not happen automatically in an error handler.

Classify outcomes. A completed operation returns its result. A running operation returns pending or lets one worker claim it. A retryable operation schedules a bounded attempt. A parameter mismatch is a conflict. A terminal rejection requires new intent. This makes retries deterministic rather than hopeful.

Add jitter and attempt budgets to protect dependencies during incidents. After the budget is exhausted, preserve the operation as blocked with evidence for an operator.

Make completion idempotent across every side effect

The final transition often creates more trouble than the main command because it fans out. Completing delivery can release payout eligibility, award loyalty credit, update analytics, notify participants, and close operational tasks. Re-running completion must not duplicate any of them.

Create a completion record or effect rows in the same transaction as the authoritative state transition. Give each effect a unique key such as operation ID plus effect type and recipient. Workers deliver effects independently and mark their own result. A notification outage does not roll back completed delivery, and a worker retry cannot create another payout operation.

Separate derived projections from irreversible effects. Search indexes, dashboards, and aggregates can be rebuilt from authoritative records. Payments and external messages need explicit identity and receipts. Both should be observable, but their recovery strategy differs.

For sockets, assume clients miss or duplicate events. Emit a versioned state change and let reconnecting clients fetch the authoritative snapshot. The socket message accelerates freshness; it is not the only ledger. For timers, store the deadline and current lifecycle generation durably. A delayed timer verifies both before transitioning so an old timer cannot finish a newly restarted game.

Completion should be safe to call whenever state is uncertain. That property simplifies webhooks, recovery jobs, and operator tools.

Test interleavings instead of only endpoints

Ordinary endpoint tests usually execute commands sequentially. Add tests that start competing operations together and assert the invariant, not which request wins. Submit accept and cancel concurrently. Run two finalizers for one game. Deliver the same webhook through two workers. Let a timer and live event race. Confirm exactly one valid terminal state and one set of effects.

Use barriers in tests to pause operations after reading but before writing, which forces the vulnerable interleaving. Inject a transaction abort, lock timeout, process crash after commit, queue redelivery, and lease expiry. Then run recovery and assert convergence.

Inspect the database after tests. Response codes alone can hide duplicate effect rows or invalid references. Assert unique operation records, monotonically increasing versions, allowed state transitions, outbox identity, and audit history. Property-based tests can generate transition sequences and verify that terminal invariants always hold.

Observe the same behavior in production through conflict rate, constraint violations, transaction retries, lock wait, deadlocks, expired leases, duplicate receipts, completion lag, and blocked operations. A rising conflict rate may mean normal contention or a missing product rule; either way it deserves visibility.

Prefer explicit conflict over silent corruption

Concurrency-safe systems do not eliminate conflict. They detect it at the boundary where truth can still be preserved. A stale client receives a version conflict. One insert loses a unique race. A transaction aborts and retries. A worker fails to acquire a lease. These are controlled outcomes, not server failures to hide.

Design APIs and interfaces to explain them. Reload current state, show who or what changed it when appropriate, and offer only safe next actions. An operator recovery panel should reuse the same transition service, operation identity, and audit trail as automation. Direct field edits bypass the very invariants the system relies on.

The core method is consistent across relational databases, document stores, queues, sockets, and providers: state the invariant, place the guard in the atomic write, constrain duplicate identity, keep critical sections short, make long ownership expire safely, and make completion repeatable. When those properties are designed together, marketplace and real-time workflows can tolerate simultaneous users and unreliable delivery without producing impossible product states.

Primary sources

  1. 1.PostgreSQL Explicit Locking — PostgreSQL
  2. 2.MongoDB Transactions — MongoDB
  3. 3.Distributed Locks with Redis — Redis
  4. 4.Stripe Idempotent Requests — Stripe

Portfolio evidence