Skip to content

Idempotency keys for email: preventing duplicate sends

A client retry after a timeout should never email a customer twice. How idempotency keys work, the race people miss and how to pick key values.

Koltrix Team4 min read
A computer screen filled with source code
Photo by Chris Ried on Unsplash
On this page(10 sections)
  1. The problem in one diagram
  2. How idempotency keys work
  3. The race condition most implementations miss
  4. Choosing good key values
  5. Should the payload be part of the key?
  6. What the client should do with each response
  7. How long to keep keys
  8. Idempotency inside your own pipeline
  9. Checklist
  10. Key takeaways

A customer clicks "Pay," your server calls the email API to send a receipt, and the connection times out. Did the send happen?

Your code cannot know, so it retries, and the customer gets two receipts. Idempotency keys exist to make that retry harmless.

The problem in one diagram

client ──POST /emails──▶ API ── queues message ──▶ (sent)
client ◀──── timeout ────X   (response lost)
client ──POST /emails──▶ API ── queues message ──▶ (sent again)

The first request succeeded, but the response never arrived. From the client's side, a timeout looks exactly like a failure. Retrying is the right instinct; without idempotency, it is also how duplicates happen.

For most API calls, a duplicate is an annoyance. For email, it is visible to your customer, it erodes trust in billing and security mail, and at scale it inflates volume and complaints.

How idempotency keys work

The client generates a unique value for each logical operation and sends it in a header:

POST /v1/emails
Idempotency-Key: receipt-order-48213

The server's contract is:

  1. First request with a new key: perform the operation, store the response against the key, return it.
  2. Repeat with the same key after completion: do not perform the operation again; return the stored response.
  3. Repeat while the first is still running: do not start a second operation; tell the client to wait and retry.

The IETF HTTP API working group has a draft specification for the Idempotency-Key header field, and many payment and messaging APIs follow this pattern. As one example, Koltrix's send endpoint accepts an Idempotency-Key and keeps the stored response for 24 hours.

The race condition most implementations miss

The naive implementation checks the key at the start and saves it at the end:

# BROKEN: check-then-act race
cached = store.get(key)
if cached:
    return cached
response = send_email(payload)      # takes time: DB writes, enqueue
store.set(key, response, ttl=86400)
return response

Between get and set, a second request with the same key, often a client retry fired because the first one was slow, sees nothing cached and sends again. This window is precisely where retries land, so the bug shows up in exactly the situation the feature was meant to handle.

The fix is to claim the key atomically before doing any work:

IN_PROGRESS = "__in_progress__"

def handle(key, payload):
    if store.set_if_absent(key, IN_PROGRESS, ttl=86400):   # atomic, e.g. Redis SET NX
        try:
            response = send_email(payload)
        except Exception:
            store.delete(key)            # nothing was sent; allow an immediate retry
            raise
        store.set(key, serialize(response), ttl=86400)
        return response, 202

    cached = store.get(key)
    if cached is None or cached == IN_PROGRESS:
        return {"error": "request with this key is still in progress"}, 409
    return deserialize(cached), 200

Three details in that sketch matter:

  • The claim is atomic. Only one request can win set_if_absent.
  • Failures before anything is queued release the key, so the client is not locked out for the whole retention window by an attempt that sent nothing. Be careful to release only when you are sure nothing was queued.
  • If the store is unreachable, fail closed. Return a 503 and send nothing. Sending without the ability to deduplicate risks exactly the duplicate you are trying to prevent, and the client can safely retry with the same key once the store is back.

Choosing good key values

The key identifies the logical operation, not the HTTP request. Good keys are deterministic for a given business event:

Event Good key Bad key
Receipt for an order receipt:order:48213 A random UUID generated per attempt
Password reset request reset:user:991:request:7a2f (request ID) reset:user:991 (blocks legitimate second reset)
Weekly digest digest:user:991:2026-W40 digest:user:991

A random UUID is fine if you generate it once per logical operation and store it with the job, so every retry of that job reuses it. Generating a fresh UUID inside the retry loop defeats the purpose.

Scope keys per account on the server side, so two customers who happen to choose the same key string do not collide.

Should the payload be part of the key?

If a client reuses a key with a different payload, something is wrong on their side. Some APIs store a hash of the request body with the key and return an error, such as 422, when a reused key arrives with a different body. That catches bugs where a key generator produces collisions. It is a worthwhile addition if you can afford the extra storage.

What the client should do with each response

Idempotency only helps if the client handles the responses sensibly. A 202 or 200 means the send is recorded; store the returned message ID with your business record. A "still in progress" response such as 409 means another attempt with the same key is running; wait a few seconds and retry with the same key rather than generating a new one. A 503 from a server that fails closed means nothing was sent; retry later with the same key. Validation errors mean the request itself is wrong; do not retry until the payload is fixed.

How long to keep keys

Long enough to cover realistic retry windows. Client retries usually happen within seconds to minutes; background job retries can span hours. Twenty-four hours is a common retention, balancing protection against storage. Document the window so callers know a retry after it may send again.

Idempotency inside your own pipeline

The API key protects the boundary between client and server. Inside your system, the same message might be processed twice by a worker that crashes after sending but before acknowledging the queue job. Defend that boundary too:

  • Record a per-message state transition, such as queued to sending to sent, with a conditional update so only one worker can move a message forward.
  • Make webhook consumers idempotent by event ID, since webhooks are typically delivered at least once.

Checklist

  • Clients generate one key per logical send and reuse it on every retry.
  • The server claims the key atomically before any work.
  • In-flight duplicates get a "still in progress" response, such as 409.
  • Completed duplicates get the original response, ideally marked as a replay.
  • Failures before queuing release the key; store outages fail closed with 503.
  • Keys are scoped per account and retained for a documented window.

Key takeaways

  • Timeouts make clients retry, and without idempotency every retry can send a duplicate email.
  • An idempotency key turns a repeated request into a replay of the original response.
  • Claim the key atomically before doing work, or concurrent retries will still duplicate.
  • Derive keys from the business event, not from the attempt.

Start with Koltrix

Your domain, one inbox, and an API that sends.

A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

SharePost on XLinkedIn