Skip to content

Retry policies for outbound mail: backoff and give-up times

How long should a sender keep retrying a deferred message? Exponential backoff with jitter, sensible give-up windows and what to tell users meanwhile.

Koltrix Team4 min read
Layered mountain ridges under a warm sky
Photo by Sven Pieren on Unsplash
On this page(9 sections)
  1. Where retries happen
  2. Classify before you retry
  3. Exponential backoff with jitter
  4. How long to keep trying
  5. Respect the receiver's signals
  6. Tell users what is happening
  7. Avoid retry amplification
  8. Policy checklist
  9. Key takeaways

Every sender has to decide what to do when a receiving server says "not now."

Retry too aggressively and you look like a spammer and get throttled harder. Give up too early and real mail never arrives. A good retry policy sits between those extremes and changes depending on what kind of message it is carrying.

Where retries happen

There are two different retry loops in most email pipelines, and they need different policies:

  1. Your application to your provider or relay. The API call or SMTP submission fails because of a timeout, a 5xx, a 429 or a network error. The message has not been accepted by your sending infrastructure yet.
  2. Your mail server or provider to the recipient's server. The recipient's MX returns a 4xx temporary failure: greylisting, rate limiting, a full mailbox, or a server problem.

If you use a sending provider, it handles the second loop for you, with its own policy. You still own the first. If you run your own MTA, you own both.

Classify before you retry

Retrying only makes sense for failures that might succeed later.

Response Meaning Retry?
Network error, timeout Unknown outcome Yes, with the same idempotency key
HTTP 429 Rate limited Yes, after Retry-After if provided
HTTP 5xx Provider-side problem Yes
HTTP 4xx other than 429 Your request is wrong No; fix and alert
SMTP 4xx (e.g. 421, 450, 451, 452) Temporary failure Yes
SMTP 5xx (e.g. 550, 553, 554) Permanent failure No; generate a bounce

Under RFC 5321, a reply code starting with 4 is a transient failure and the sender is expected to retry; a code starting with 5 is permanent. Enhanced status codes (RFC 3463) add detail but follow the same class. Some receivers use 4xx codes for situations that will never resolve, and occasionally 5xx for things that would; that is why a retry policy also needs a maximum duration.

Exponential backoff with jitter

The standard pattern is exponential backoff: each retry waits longer than the last. Jitter adds randomness so many messages that failed at the same moment do not all retry at the same moment.

import random

def next_delay(attempt: int, base: float = 30.0, cap: float = 3600.0) -> float:
    """Full-jitter exponential backoff, in seconds."""
    exp = min(cap, base * (2 ** attempt))
    return random.uniform(0, exp)

"Full jitter," choosing uniformly between zero and the exponential ceiling, spreads load well during recovery from an outage. Some systems prefer a gentler variant that keeps a minimum delay, such as half the ceiling plus a random half, so retries never fire immediately.

Without jitter, a provider outage that fails 100,000 messages at 10:00 produces 100,000 retries at 10:01, 100,000 at 10:03 and so on, hitting the provider exactly as it recovers.

How long to keep trying

RFC 5321 suggests that SMTP retry intervals should generally be at least 30 minutes and that senders should keep trying for at least four to five days before giving up. That guidance reflects the protocol's store-and-forward heritage, and it remains a sensible default for ordinary mail.

But for many transactional messages, late delivery is worse than no delivery. Match the give-up time to the message's useful life:

Message type Useful life Suggested give-up
Sign-in code, OTP Minutes Shortly after the code expires
Password reset link Under an hour When the link expires
Security alert Hours 24 hours
Receipt or invoice Days Several days, per RFC guidance
Digest or newsletter Until the next one Before the next edition

Attach an expiry to each message when it is created, and have workers drop messages past their expiry instead of delivering them. A sign-in code that arrives two hours late only confuses the user.

Respect the receiver's signals

Large mailbox providers use temporary failures to slow senders down. A 421 response with text like "try again later" or a reference to rate limits is a request to reduce throughput, not just to retry one message.

  • Back off per destination. If one provider is deferring you, slow all traffic to that provider, not just the deferred message.
  • Reduce concurrency to that destination: fewer simultaneous connections and fewer messages per connection.
  • Avoid retrying from many IPs at once, which can look like evasion and may defeat greylisting.
  • Increase gradually once deferrals subside.

This adaptive behavior matters more than the exact backoff formula.

Tell users what is happening

When a message is in retry, the application should know. Expose a status such as "delayed, retrying" so support can answer questions, and consider product affordances for time-sensitive mail: a "resend code" button that issues a fresh code rather than retrying a stale one, or a note that delivery may take a few minutes.

When a message finally gives up, record why. "Gave up after 5 days; last response 452 4.2.2 mailbox full" is actionable. "Failed" is not.

Avoid retry amplification

Retries multiply load. If your application retries five times, your queue library retries three times per attempt, and your HTTP client retries twice per call, a single failure can generate thirty requests. Pick one layer to own retries, usually the queue worker, and set the others to fail fast.

Similarly, always retry with the same idempotency key, so a request that actually succeeded before timing out is not sent again.

Policy checklist

  • Classify responses: retry 4xx SMTP, 429, 5xx HTTP and network errors; never retry permanent failures.
  • Use exponential backoff with jitter.
  • Honor Retry-After.
  • Set give-up times per message type based on useful life.
  • Back off per destination when a receiver defers broadly.
  • Retry in one layer only, with the same idempotency key.
  • Record the final response and reason for every give-up.

Key takeaways

  • There are two retry loops, app-to-provider and provider-to-recipient, and each needs a policy.
  • Retry temporary failures with exponential backoff and jitter; never retry permanent ones.
  • RFC 5321's multi-day retry guidance suits ordinary mail, but time-sensitive messages should expire much sooner.
  • Deferrals from big providers are a signal to slow down overall, not just to retry.

Start with Koltrix

Your domain, one inbox, and an API that sends.

A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

SharePost on XLinkedIn