Observability for email pipelines: logs, metrics, traces
Trace a message from API call to SMTP response. What to log, which metrics to emit and how to correlate IDs so 'did it send?' takes seconds.

On this page(9 sections)
"Did it send?" should take ten seconds to answer, not an afternoon of grepping logs across four services. Email pipelines are distributed systems with an external dependency at the end, and they deserve the same observability as any other critical path: correlated logs, a handful of meaningful metrics, and traces that follow a message from API call to SMTP response.
The path you need to see
A typical transactional send crosses several boundaries:
- Trigger: a user action or a scheduled job decides an email is needed.
- Enqueue: the application writes a job to a queue.
- Render: a worker renders the template with data.
- Submit: the worker calls the email provider's API or SMTP relay.
- Accept: the provider acknowledges and queues the message.
- Deliver: the provider's servers hand the message to the recipient's mail server, which accepts, defers or rejects it.
- Events: webhooks report delivery, bounces, complaints and engagement, sometimes hours later.
Problems can occur at every step, and the symptom ("the customer didn't get it") looks the same regardless of where it happened. Observability means being able to tell which step failed.
Correlation IDs: the foundation
Pick one identifier for each logical message and carry it everywhere:
- Generate it when the email is first requested, before enqueueing.
- Include it in every log line about that message.
- Store it in your messages table.
- Pass it to the provider, as the idempotency key, a custom header, or metadata, if the provider supports any of those.
- Record the provider's own message ID when the API responds, and map between the two.
When a webhook arrives later carrying the provider's ID, you can join it back to your ID and see the whole story.
Also keep the business context: which user, which account, which invoice or ticket. Support staff search by customer, not by UUID.
Structured logs
Log each step as a structured event rather than free text:
{
"ts": "2026-09-30T14:02:11.482Z",
"level": "info",
"event": "email.submitted",
"email_id": "6f1c2b0e-8d2a-4a51-9a7e-1f0c9b7d2e44",
"template": "invoice-paid",
"template_version": "2026-09-12",
"account_id": "acct_1842",
"recipient_domain": "example.net",
"provider": "primary",
"provider_message_id": "msg_01J9Z...",
"status_code": 202,
"latency_ms": 184
}
Useful event names follow the path: email.requested, email.enqueued, email.rendered, email.submitted, email.provider_rejected, email.delivered, email.bounced, email.complained.
Be deliberate about personal data. Logging the recipient's domain is usually enough for operational analysis; full addresses and message bodies in general-purpose logs create privacy and retention problems. Keep full details in the email events table, where access and retention are controlled.
Metrics worth emitting
A small set of metrics covers most operational questions:
| Metric | Type | Why |
|---|---|---|
| Emails requested, by template | Counter | Detects missing or duplicated triggers |
| Queue depth and oldest job age | Gauge | Detects stuck or slow workers |
| Render failures, by template | Counter | Catches template bugs after deploys |
| Provider submit latency | Histogram | Detects provider slowness |
| Provider responses, by status code | Counter | Separates 2xx, 4xx, 429 and 5xx |
| Delivered, bounced, complained, by template and recipient domain | Counter | Delivery health |
| Time from request to delivered event | Histogram | End-to-end user experience |
The end-to-end latency metric is the one customers feel. A sign-in code that takes four minutes to arrive is effectively a failure, even if every component reports success.
Watch label cardinality: label by template, provider and recipient domain group, never by recipient address or message ID.
Traces across the queue
Distributed tracing shows a request's journey through services. Email adds a twist: the queue breaks the synchronous chain. To keep the trace connected, propagate trace context through the job payload:
- When enqueueing, attach the current trace context (for example, W3C Trace Context headers) to the job.
- When the worker picks up the job, start a span linked to that context.
- Wrap rendering and the provider call in child spans, with the email ID as an attribute.
Now a single trace shows the HTTP request that triggered the email, the time spent waiting in the queue, the render, and the provider call. Queue wait time often turns out to be the biggest contributor to slow delivery.
Webhook events arrive too late to be part of the original trace. Link them through the email ID instead.
Alerts that matter
Alert on symptoms, with thresholds based on your normal traffic:
- Oldest queued job older than a few minutes for high-priority templates such as sign-in codes and password resets.
- Provider error rate (5xx and timeouts) above baseline.
- 429 rate rising, indicating you are hitting rate limits.
- Render failures greater than zero for any template after a deploy.
- Bounce or complaint rate for a template well above its baseline.
- Delivered events stop arriving while submissions continue, which may mean the webhook pipeline is broken rather than delivery.
That last one is easy to miss. If your webhook endpoint starts failing, your dashboards may show emails as stuck in "sent" while customers receive them fine, or worse, hide real bounces.
A support lookup in one place
Give support a simple view keyed by customer or email address that shows each message, its template, timestamps for each step, the provider response, delivery events and bounce reasons. Most "did it send?" tickets can then be answered without involving engineering:
- Never requested: a product or trigger bug.
- Requested but not submitted: a queue or render problem.
- Submitted and delivered: likely filtered or overlooked by the recipient; suggest checking spam.
- Bounced: show the reason and ask for a corrected address.
Checklist
- One email ID generated at request time and carried through every step.
- Provider message IDs mapped to your IDs.
- Structured logs for each pipeline step, without full addresses or bodies.
- Metrics for queue age, provider responses, delivery outcomes and end-to-end latency.
- Trace context propagated through the queue.
- Alerts on queue age, provider errors, render failures, bounce spikes and missing webhook events.
- A support view that answers "did it send?" per customer.
Key takeaways
- Email delivery spans your systems and a provider's, so correlation IDs are essential.
- Structured logs, a small metric set and queue-aware traces show where a message got stuck.
- End-to-end time from request to delivery is the metric users actually experience.
- Alert on missing webhook events as well as on errors.
- A good support view turns most email tickets into a self-service lookup.
Start with Koltrix
Your domain, one inbox, and an API that sends.
A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.


