Five deliverability metrics worth alerting on
Skip vanity numbers. Alert on deferral rate, hard bounce rate, complaint rate, authentication failures and delivery latency, with sensible thresholds.

On this page(11 sections)
Most email dashboards are built to make marketing look good, not to wake engineers when delivery breaks. If you are responsible for mail actually arriving, you need a short list of metrics that move early, mean something specific, and come with a clear first action.
What makes a metric worth alerting on
A good deliverability alert has three properties:
- It changes before customers notice. By the time support tickets say "I never got the email," you are late.
- It points to a cause. "Opens are down" could mean anything. "Deferrals from one provider tripled" tells you where to look.
- It is computed per stream and per receiving provider. A problem at one mailbox provider gets lost in a global average.
With that in mind, these five metrics cover most real incidents.
1. Deferral rate
What it is: the share of delivery attempts that receive a temporary 4xx response, per receiving provider.
Why it matters: deferrals are how receivers say "slow down" or "we are not sure about you." They rise before outright blocks and spam-folder placement, which makes them the best early warning you have. They also directly delay mail: a password reset deferred for 20 minutes is effectively broken.
How to compute: for each provider group (Gmail, Outlook.com, Yahoo and "everything else" is a sensible start), divide 4xx responses by total delivery attempts over a rolling window.
Alert when: the rate for a provider rises well above its usual baseline for that time of day, for example doubling over an hour, or stays elevated for a sustained period.
First action: read the deferral text. Rate-limit messages point to throttling and volume; messages mentioning reputation or policy point to a sender problem.
2. Hard bounce rate
What it is: permanent failures (5xx, typically "user unknown" or "mailbox does not exist") divided by messages sent, per stream.
Why it matters: a sudden increase usually means bad addresses entered the system: a list import, a broken signup form, a bot attack creating accounts, or a bug sending to the wrong field. High bounce rates also damage reputation directly.
Alert when: the rate for a stream jumps above its normal level. For transactional mail to verified users it should be very low; for mail to newly collected addresses it will be higher.
First action: group bounces by source (which signup path, which import, which feature) and by recipient domain.
3. Spam complaint rate
What it is: complaints divided by delivered messages, per stream. Sources are feedback loop reports where providers offer them, and aggregate rates in Google Postmaster Tools for Gmail, which does not send per-message reports.
Why it matters: Gmail and Yahoo both tell bulk senders to keep spam rates below 0.3%, and Gmail recommends staying under 0.1%. Complaints are among the strongest negative signals a receiver can get.
Alert when: the rate for a stream approaches 0.1%, or rises sharply after a specific campaign.
First action: identify the campaign or message type, then check targeting, frequency and how easy it is to unsubscribe.
4. Authentication failure rate
What it is: the share of your own outbound messages that fail SPF, DKIM or DMARC at receivers, from DMARC aggregate reports and your own test probes.
Why it matters: an expired key, a DNS change that removed an include, or a new system sending without DKIM can silently break authentication. At bulk volumes, failing DMARC now leads to rejections at major providers.
Alert when: DMARC aggregate data shows a drop in the aligned pass rate for one of your known sources, or a scheduled probe message fails authentication.
First action: identify the source IP or selector that started failing and check recent DNS and configuration changes.
5. Delivery latency
What it is: time from when your application requests a send to when the receiving server accepts the message (a 250 response), measured as percentiles.
Why it matters: latency is what users experience. Verification codes and magic links that arrive after five minutes cause abandoned signups even if every message eventually delivers. Latency rises with queue backlogs, deferrals, provider incidents and worker problems.
Alert when: the 95th percentile for security and transactional mail exceeds a target you set, such as one minute.
First action: check whether the delay is in your queue (backlog, worker health) or at the receiver (deferrals).
A compact alert table
| Metric | Granularity | Typical source | Example trigger |
|---|---|---|---|
| Deferral rate | Per provider, per stream | MTA or provider events | 2x baseline for 30+ minutes |
| Hard bounce rate | Per stream, per source | Bounce events | Sharp jump over daily norm |
| Complaint rate | Per stream, per campaign | FBL reports, Postmaster Tools | Approaching 0.1% |
| Auth failure rate | Per source IP or selector | DMARC reports, probes | Any known source drops below normal pass rate |
| Delivery latency p95 | Per stream | Send and delivery timestamps | Above target for 10+ minutes |
These thresholds are examples. Set yours from your own baseline after a few weeks of data.
What not to alert on
- Open rates. Privacy features in some mail clients load images automatically, which inflates opens and makes them noisy. Useful for trends; poor for paging someone.
- Total sent volume on its own. Volume changes for business reasons. Alert on unexpected changes relative to what your application requested, not on volume itself.
- Every single bounce or complaint. Individual events belong in dashboards and logs; rates belong in alerts.
Synthetic probes fill the gaps
Some failures do not show up in aggregate metrics quickly. A scheduled job that sends a real message every few minutes to test mailboxes at major providers, then checks that it arrived, how long it took and whether authentication passed, catches broken DNS, expired credentials and stuck queues fast. Keep probe volume tiny and send it from the same infrastructure and domains as production.
Checklist
- Break every metric down by stream and receiving provider.
- Alert on deferral rate, hard bounce rate, complaint rate, authentication failures and delivery latency.
- Derive thresholds from your own baselines.
- Attach a first diagnostic step to each alert.
- Add synthetic probes for end-to-end coverage.
Key takeaways
- Alert on metrics that move early and point to a cause: deferrals, hard bounces, complaints, authentication failures and latency.
- Always compute them per stream and per receiving provider.
- Gmail and Yahoo's 0.3% spam rate line, and Gmail's 0.1% recommendation, make complaint rate a hard limit, not a vanity metric.
- Open rates and raw volume are trend data, not alert material.
- Synthetic probes catch silent breakage that aggregates miss.
Start with Koltrix
Your domain, one inbox, and an API that sends.
A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

