Building webhook consumers for retries and disorder
Webhooks arrive late, twice or out of order. Design consumers that acknowledge fast, deduplicate by event, process asynchronously and tolerate disorder.

On this page(9 sections)
- What "at least once" and "unordered" really mean
- At least once
- No ordering guarantee
- Pattern 1: acknowledge fast, process later
- Pattern 2: deduplicate by event identity
- Pattern 3: make state transitions order-independent
- Use the provider's timestamp, not arrival time
- Pattern 4: tolerate events for unknown messages
- Pattern 5: reconcile periodically
- Monitoring the consumer
- Checklist
- Key takeaways
Every email provider's webhook documentation contains some version of the same sentence: events are delivered at least once, and order is not guaranteed. Most consumers are written as if neither were true, and they work fine until the day a bounce arrives before the "sent" event and a user is shown as active on a dead address.
What "at least once" and "unordered" really mean
At least once
A provider sends an event and waits for a 2xx response. If your endpoint times out, returns an error, or the response is lost on the network, the provider cannot know whether you processed the event, so it sends it again later. You will therefore occasionally receive the same event twice or more. Exactly-once delivery across a network is not something any provider can promise.
No ordering guarantee
Events for one message are generated by different parts of the provider's system at different moments, delivered over separate HTTP requests, retried independently, and possibly processed by different workers on your side. A message's bounced event can arrive before its sent event if the first delivery of sent failed and was retried. An opened event can arrive after a complained event.
Design for both from the start. Retrofitting is painful.
Pattern 1: acknowledge fast, process later
The most important structural decision is to separate receiving from processing.
provider ──▶ endpoint: verify signature ─▶ insert raw event ─▶ 200 OK
│
▼
worker: dedupe ─▶ apply to state ─▶ side effects
The endpoint does only three things: verify the signature, persist the raw event durably, and return 200. Everything else happens in a background worker.
Why this matters:
- Provider timeouts are short. Many providers wait only a few seconds; Koltrix, for example, attempts each webhook once with an 8-second timeout. With providers that retry, a slow handler causes duplicates; with providers that do not, it causes missed events.
- Your downstream systems will sometimes be slow or down. If processing fails, the raw event is already stored and can be retried on your schedule.
- You keep an audit trail. The raw events table answers "what did the provider tell us, and when?"
Pattern 2: deduplicate by event identity
Store raw events with a unique constraint on the provider's event ID, if one is supplied. If not, derive a key from fields that together identify the event: message ID, event type and the provider's event timestamp.
CREATE TABLE webhook_events (
event_key text PRIMARY KEY,
message_id text NOT NULL,
type text NOT NULL,
occurred_at timestamptz NOT NULL,
payload jsonb NOT NULL,
received_at timestamptz NOT NULL DEFAULT now(),
processed_at timestamptz
);
INSERT INTO webhook_events (event_key, message_id, type, occurred_at, payload)
VALUES ($1, $2, $3, $4, $5)
ON CONFLICT (event_key) DO NOTHING;
A duplicate delivery becomes a no-op insert, and the endpoint still returns 200 so the provider stops retrying.
Side effects need the same protection. If processing an event sends a Slack alert or updates a CRM, make that step idempotent too, for example by recording that the side effect was performed for this event key.
Pattern 3: make state transitions order-independent
Instead of applying events as commands ("set status to sent"), treat each message's status as derived from all events received so far, with explicit precedence.
| Status | Precedence | Notes |
|---|---|---|
| complained | highest | Recipient reported spam; drives suppression |
| bounced | high | Permanent failure; drives suppression |
| unsubscribed | high | Category-level suppression |
| clicked | medium | Implies opened and sent |
| opened | medium | Implies sent; unreliable with privacy proxies |
| sent | low | Accepted by the receiving server |
| queued | lowest | Accepted by the provider |
When an event arrives, update the message's status only if the new event's precedence is higher than the current one:
UPDATE messages
SET status = $2, status_rank = $3, status_at = $4
WHERE id = $1 AND status_rank < $3;
A late sent event after bounced is recorded in the events table but does not overwrite the bounce. A late opened after clicked changes nothing.
For counters, such as open counts, count distinct events rather than incrementing on each delivery, or deduplication protects you automatically.
Use the provider's timestamp, not arrival time
When you need chronology, such as "first opened at," use the event time in the payload, not the time you received it. Retries can delay arrival by minutes or hours.
Pattern 4: tolerate events for unknown messages
Occasionally an event arrives for a message ID your database has not stored yet, for example when the provider's webhook beats your own write after a send call. Do not reject it with a 4xx; that triggers retries and possibly disables your endpoint. Store it, and let the worker retry processing later or reconcile when the message record appears.
Pattern 5: reconcile periodically
Webhooks can be missed entirely: your endpoint was down longer than the provider's retry window, or a deploy broke signature verification for a day. Many providers retry for a limited period, after which the delivery is marked failed.
Run a reconciliation job that, for recent messages without a final status, queries the provider's API for the current state. That turns occasional webhook loss from silent data corruption into a delayed but correct result.
Monitoring the consumer
- Endpoint error rate and latency. Rising latency predicts retries.
- Signature failures. A spike usually means a rotated secret or a parsing change.
- Worker backlog. Unprocessed raw events growing over time.
- Duplicate rate. Some is normal; a sudden increase means your endpoint is slow or failing.
- Messages stuck without a final status beyond the expected window.
Checklist
- Endpoint verifies, stores the raw event and returns
2xxwithin a second or two. - Raw events are deduplicated with a unique event key.
- Side effects are idempotent per event key.
- Status is derived with precedence rules, not overwritten by arrival order.
- Provider timestamps are used for chronology.
- Events for unknown messages are stored, not rejected.
- A reconciliation job covers missed deliveries.
Key takeaways
- Webhooks arrive late, more than once and out of order; design for all three.
- Acknowledge quickly and process asynchronously from a durable raw events table.
- Deduplicate by event identity and make side effects idempotent.
- Derive message status from precedence rules so late events cannot regress it.
- Reconcile against the provider's API to catch deliveries that never arrived.
Start with Koltrix
Your domain, one inbox, and an API that sends.
A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.


