Skip to content

Building webhook consumers for retries and disorder

Webhooks arrive late, twice or out of order. Design consumers that acknowledge fast, deduplicate by event, process asynchronously and tolerate disorder.

Koltrix Team4 min read
Yellow and green patch cables neatly connected to a panel
Photo by Albert Stoynov on Unsplash
On this page(9 sections)
  1. What "at least once" and "unordered" really mean
  2. At least once
  3. No ordering guarantee
  4. Pattern 1: acknowledge fast, process later
  5. Pattern 2: deduplicate by event identity
  6. Pattern 3: make state transitions order-independent
  7. Use the provider's timestamp, not arrival time
  8. Pattern 4: tolerate events for unknown messages
  9. Pattern 5: reconcile periodically
  10. Monitoring the consumer
  11. Checklist
  12. Key takeaways

Every email provider's webhook documentation contains some version of the same sentence: events are delivered at least once, and order is not guaranteed. Most consumers are written as if neither were true, and they work fine until the day a bounce arrives before the "sent" event and a user is shown as active on a dead address.

What "at least once" and "unordered" really mean

At least once

A provider sends an event and waits for a 2xx response. If your endpoint times out, returns an error, or the response is lost on the network, the provider cannot know whether you processed the event, so it sends it again later. You will therefore occasionally receive the same event twice or more. Exactly-once delivery across a network is not something any provider can promise.

No ordering guarantee

Events for one message are generated by different parts of the provider's system at different moments, delivered over separate HTTP requests, retried independently, and possibly processed by different workers on your side. A message's bounced event can arrive before its sent event if the first delivery of sent failed and was retried. An opened event can arrive after a complained event.

Design for both from the start. Retrofitting is painful.

Pattern 1: acknowledge fast, process later

The most important structural decision is to separate receiving from processing.

provider ──▶ endpoint: verify signature ─▶ insert raw event ─▶ 200 OK
                                              │
                                              ▼
                              worker: dedupe ─▶ apply to state ─▶ side effects

The endpoint does only three things: verify the signature, persist the raw event durably, and return 200. Everything else happens in a background worker.

Why this matters:

  • Provider timeouts are short. Many providers wait only a few seconds; Koltrix, for example, attempts each webhook once with an 8-second timeout. With providers that retry, a slow handler causes duplicates; with providers that do not, it causes missed events.
  • Your downstream systems will sometimes be slow or down. If processing fails, the raw event is already stored and can be retried on your schedule.
  • You keep an audit trail. The raw events table answers "what did the provider tell us, and when?"

Pattern 2: deduplicate by event identity

Store raw events with a unique constraint on the provider's event ID, if one is supplied. If not, derive a key from fields that together identify the event: message ID, event type and the provider's event timestamp.

CREATE TABLE webhook_events (
  event_key   text PRIMARY KEY,
  message_id  text NOT NULL,
  type        text NOT NULL,
  occurred_at timestamptz NOT NULL,
  payload     jsonb NOT NULL,
  received_at timestamptz NOT NULL DEFAULT now(),
  processed_at timestamptz
);

INSERT INTO webhook_events (event_key, message_id, type, occurred_at, payload)
VALUES ($1, $2, $3, $4, $5)
ON CONFLICT (event_key) DO NOTHING;

A duplicate delivery becomes a no-op insert, and the endpoint still returns 200 so the provider stops retrying.

Side effects need the same protection. If processing an event sends a Slack alert or updates a CRM, make that step idempotent too, for example by recording that the side effect was performed for this event key.

Pattern 3: make state transitions order-independent

Instead of applying events as commands ("set status to sent"), treat each message's status as derived from all events received so far, with explicit precedence.

Status Precedence Notes
complained highest Recipient reported spam; drives suppression
bounced high Permanent failure; drives suppression
unsubscribed high Category-level suppression
clicked medium Implies opened and sent
opened medium Implies sent; unreliable with privacy proxies
sent low Accepted by the receiving server
queued lowest Accepted by the provider

When an event arrives, update the message's status only if the new event's precedence is higher than the current one:

UPDATE messages
SET status = $2, status_rank = $3, status_at = $4
WHERE id = $1 AND status_rank < $3;

A late sent event after bounced is recorded in the events table but does not overwrite the bounce. A late opened after clicked changes nothing.

For counters, such as open counts, count distinct events rather than incrementing on each delivery, or deduplication protects you automatically.

Use the provider's timestamp, not arrival time

When you need chronology, such as "first opened at," use the event time in the payload, not the time you received it. Retries can delay arrival by minutes or hours.

Pattern 4: tolerate events for unknown messages

Occasionally an event arrives for a message ID your database has not stored yet, for example when the provider's webhook beats your own write after a send call. Do not reject it with a 4xx; that triggers retries and possibly disables your endpoint. Store it, and let the worker retry processing later or reconcile when the message record appears.

Pattern 5: reconcile periodically

Webhooks can be missed entirely: your endpoint was down longer than the provider's retry window, or a deploy broke signature verification for a day. Many providers retry for a limited period, after which the delivery is marked failed.

Run a reconciliation job that, for recent messages without a final status, queries the provider's API for the current state. That turns occasional webhook loss from silent data corruption into a delayed but correct result.

Monitoring the consumer

  • Endpoint error rate and latency. Rising latency predicts retries.
  • Signature failures. A spike usually means a rotated secret or a parsing change.
  • Worker backlog. Unprocessed raw events growing over time.
  • Duplicate rate. Some is normal; a sudden increase means your endpoint is slow or failing.
  • Messages stuck without a final status beyond the expected window.

Checklist

  • Endpoint verifies, stores the raw event and returns 2xx within a second or two.
  • Raw events are deduplicated with a unique event key.
  • Side effects are idempotent per event key.
  • Status is derived with precedence rules, not overwritten by arrival order.
  • Provider timestamps are used for chronology.
  • Events for unknown messages are stored, not rejected.
  • A reconciliation job covers missed deliveries.

Key takeaways

  • Webhooks arrive late, more than once and out of order; design for all three.
  • Acknowledge quickly and process asynchronously from a durable raw events table.
  • Deduplicate by event identity and make side effects idempotent.
  • Derive message status from precedence rules so late events cannot regress it.
  • Reconcile against the provider's API to catch deliveries that never arrived.

Start with Koltrix

Your domain, one inbox, and an API that sends.

A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

SharePost on XLinkedIn