All
Fav
0%

Idempotency

Retries hide transient failures, but a retry of a write is risky. A user pays $49.99 for an order. The app sends POST /payments and waits. The request times out. The app retries, and the second attempt succeeds. Did the card get charged once or twice?

Why a Timeout Can't Be Trusted

A timeout tells the client only that no response arrived in time. It does not say where the failure happened. There are three places it could have happened.

One timeout, three possible causes

In the first case, the network dropped the request, so the server never saw it. In the second case, the server crashed before it committed anything. In the third case, the server charged the card and committed the payment, and then the response was lost on the way back. The customer paid in case three but not in cases one and two.

The client sees the same timeout in all three cases. If it does not retry, cases one and two lose a sale the customer wanted. If it retries, cases one and two work, but case three charges the customer a second time.

Not retrying is also not under your control. Users tap Pay again. HTTP client libraries and proxies retry failed requests. A message queue redelivers a message when the consumer does not acknowledge it. The goal is therefore a retry that cannot repeat the effect.

What Idempotent Means

An operation is idempotent if running it twice leaves the system in the same state as running it once.

Consider an account with $90. One request says "add $10 to the balance". Another says "set the balance to $100". Run each one twice, as a retry would. The add ends at $110, which is $10 too much. The set ends at $100 both times, because the second run writes a value that is already there.

Many operations are naturally idempotent. A read changes nothing. A PUT that replaces a whole record writes the same record again. A DELETE of a deleted row has nothing left to delete. An upsert (insert the row, or overwrite it if a row with that key exists) is a common way to turn a write into a "set".

A charge cannot become a "set". Every run moves money. The fix is to give the request an identity that the server can recognize when it arrives again.

Idempotency Keys

An idempotency key is a unique value that the client generates for one logical operation, such as one purchase. A random UUID is typical. The client sends the key with the request, usually in a header, and sends the same key on every retry.

The server keeps a table of keys it has handled, with the response it returned for each key. On each request, the server looks up the key first. A new key means the server does the work and saves the result under the key. A known key means the server skips the work and returns the saved result.

Idempotency key: three attempts, one charge

Suppose two responses are lost and the third arrives. Attempt 1 charges $49.99 and saves 9f2c → pay_1. Attempts 2 and 3 find 9f2c and return pay_1 without a charge. The client made three attempts and the card was charged once. Without the key, the same three attempts charge $149.97.

The client must create the key once, before the first attempt. A common bug builds the whole request inside the retry loop, which generates a new key on every attempt. The server then sees three different purchases and charges three times. Nothing raises an error.

# Wrong: a new key on every attempt for attempt in range(3): key = uuid4() send("POST /payments", key, amount) # Right: one key per purchase, reused on every retry key = uuid4() for attempt in range(3): send("POST /payments", key, amount)

Stripe's API works this way. It saves the status code and body of the first request for a key, and it returns the same result for later requests with that key (Stripe docs).

The simulation below runs the same three attempts against a server with and without key checks, and against a client that makes a new key per attempt.

Idempotency Keys: Three Attempts, Lost Responses

Loading Python environment... This may take a few seconds on first load.

The second run charges once and returns pay_1 three times. The third run charges three times even though the server checks keys.

Two Retries at the Same Time

The basic key check still has gaps. The first one appears when two requests with the same key arrive together. This happens when a user double-taps Pay, or when a client retries while the first attempt is still running slowly.

Two requests with the same key at the same time

With "check, then insert", request A looks up 9f2c and finds nothing. Before A saves anything, request B looks up 9f2c and also finds nothing. Both requests charge the card. The check and the write are two steps, and the other request runs between them. This kind of bug is a race condition: the result depends on the order in which two requests interleave.

The fix reverses the order. The server first inserts the key, before any work, with a unique constraint on the key column. A unique constraint makes the database reject a second row with the same value. Request A's insert succeeds with status in_progress. Request B's insert fails because the key exists. The database picks the winner, not the application code.

Request A charges the card, then marks the key done with the result pay_1. Request B returns 409 Conflict (the request is still in progress, retry later). When the client retries B, the key is done, and the server returns pay_1.

A Crash in the Middle

The second gap is a crash. Suppose the server saves the payment in one write and marks the key done in a second write. If the process crashes between the two writes, the key stays at in_progress.

Crash between the work and the key record

The server cannot tell whether that first attempt is still running or dead. If retries wait, the client sees "in progress" forever. If a timeout lets a retry take over the key, the retry does the work again, and the customer pays twice.

The fix is to put the key and the work in the same database transaction. Inside one transaction, the server inserts the key, inserts the payment, marks the key done, and commits. A crash before the commit discards all three writes, the key included, so the retry starts fresh and charges once. A crash after the commit leaves all three writes, so the retry finds the key. The rule to state in an interview: record the key in the same transaction as the effect it protects. For more on transaction scope, see Database Transactions.

When the Work Happens in Another System

The transaction rule needs the effect to live in your database. In a real payment system, a payment provider charges the card through its own API. Your transaction cannot include their system, so your commit and their charge cannot be atomic.

When the work happens in another system

A payment worker calls the provider, and the call times out. The worker is now in the client's position from the first section: the charge may or may not have happened. Two practices handle this.

First, do not mark the payment failed. A timeout means the worker did not hear back. Put the payment into an explicit unknown state such as in_doubt instead of guessing.

Second, pass the same idempotency key down to the provider. A provider that supports keys returns the original charge when the worker retries, instead of charging again. For anything still unresolved, a reconciliation job compares your in_doubt payments against the provider's settlement report (the provider's record of what was actually charged) and finalizes each payment.

Reconciliation. A batch job matches internal records against an external source of truth. Example: payment pay_1 is in_doubt in your database, and the provider's daily report lists charge ch_1 for key 9f2c. The job marks pay_1 as captured and writes the ledger entries.

The Payment System solution builds this flow step by step, from a version that double charges to the idempotency gate, the in_doubt state, reconciliation, and a double-entry ledger.

Queues and Jobs: Exactly-Once Effects

Many systems have no client waiting for a response. Work goes through a queue, and a worker processes it. Most queues give at-least-once delivery: the worker processes a message and then sends an acknowledgment (ack). If the worker crashes before the ack, the queue delivers the message again. Every message is processed, and some are processed twice. Delivery Guarantees compares at-most-once, at-least-once, and exactly-once delivery.

Job schedulers show the same problem with leases. A lease is a claim on a job with a time limit. If the worker does not renew it, the lease expires and another worker can take the job.

A lease expires during a pause: the job runs twice

Worker A claims a job with a 30-second lease and then freezes for 45 seconds in a long garbage-collection pause. At 30 seconds the lease expires, and worker B claims and runs the job. At 45 seconds, A resumes and runs the job as well. A longer lease does not fix this, because a pause can always exceed the lease.

This is why exactly-once delivery is not practical across machines. There is always a gap between doing the work and recording that it was done. What a system can achieve is an exactly-once effect: at-least-once delivery plus idempotent work. The job carries a key, and the worker checks and records the key in one atomic step before the effect, the same insert-first move as before. Worker B records the key and runs the job. Worker A finds the key and does nothing. The Distributed Job Scheduler solution covers leases, crash recovery, and this effectively-once execution in detail.

The Same Pattern at Every Hop

Every case above has a sender that might send twice and a receiver that must not repeat the effect.

Sender → receiverWho makes the keyWhere it is checked
Client → payment APIThe client, once per purchaseAPI server, in the same transaction as the payment
Payment worker → providerYour system, passed downThe provider's idempotency layer
Queue → workerThe producer, in the message or jobThe worker, atomically before the effect

The sender keeps the key the same across retries. The receiver records the key together with the work.

Retrying Without Making Things Worse

Once retries are safe, the retry policy decides how much load they add. Retries covers the code. The rules are short.

Retry only errors that can go away: timeouts, connection errors, 429, and 5xx responses. Do not retry 400, 401, 403, or 404, because the same request returns the same error. Wait longer before each retry with exponential backoff: 1 second, then 2, 4, and 8, which is 15 seconds of waiting across four retries. Then give up.

Backoff alone leaves retries synchronized. If 1,000 clients fail at the same moment, all of them retry at exactly 1 second. Jitter adds randomness to each wait. If each client waits between 0.5 and 1 second, the 1,000 retries spread over half a second, about 200 in each 100 ms interval.

Keys also need a retention period. Keep a key longer than any client could still retry with it. Stripe, for example, may remove keys once they are at least 24 hours old, and a reused key after that starts a new request (Stripe docs).

Summary

A timeout does not say whether the work happened, so retries must be safe. The client makes one idempotency key per operation and reuses it. The server inserts the key first under a unique constraint, and it records the key in the same transaction as the effect. When the effect is in another system, pass the key down, mark timeouts in_doubt, and reconcile. Queues and job schedulers use the same pattern: at-least-once delivery plus idempotent work gives an exactly-once effect.