Skip to content

Rate limits for email APIs: handling 429s gracefully

Respect Retry-After, smooth bursts with a token bucket and keep retries idempotent. A practical guide to living within an email API's rate limits.

Koltrix Team4 min read
A traffic light hanging from a metal pole
Photo by CARTER SAUNDERS on Unsplash
On this page(10 sections)
  1. What a 429 means
  2. Kinds of limits you may hit
  3. Rule 1: respect Retry-After
  4. Rule 2: keep retries idempotent
  5. Rule 3: smooth bursts before they reach the API
  6. Rule 4: prioritize
  7. Rule 5: watch for retry storms
  8. Monitoring
  9. Checklist
  10. Key takeaways

Every email API has rate limits, whether or not its documentation makes them prominent. The day you discover yours is usually the day a backlog flushes, a bulk import triggers welcome emails, or a retry storm turns one failure into thousands of requests.

Handling HTTP 429 well is mostly about not causing it in the first place.

What a 429 means

HTTP 429 Too Many Requests, defined in RFC 6585, tells the client it has sent too many requests in a given amount of time. The response may include a Retry-After header indicating how long to wait, expressed either as a number of seconds or as an HTTP date:

HTTP/1.1 429 Too Many Requests
Retry-After: 30
Content-Type: application/json

{"error": "rate limit exceeded"}

Many APIs also return informational headers on every response, commonly named something like X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset, or the RateLimit header fields being standardized by the IETF. Names and semantics vary by provider, so read your provider's documentation rather than assuming.

The key point: a 429 means the request was not processed. Nothing was sent. Retrying is correct, as long as you wait.

Kinds of limits you may hit

Email APIs commonly enforce several limits at once:

Limit type Example What triggers it
Request rate Requests per second per API key Bursty application code, parallel workers
Message volume Messages per hour or day per account Bulk sends, plan quotas
Recipient count Recipients per request Large To/Cc/Bcc lists
Concurrency Simultaneous open requests Many workers without a pool limit
Per-domain sending Messages per period to one destination Pacing to protect deliverability

Some of these return 429; quota limits sometimes return a different error, such as 403 or a 4xx with a specific error code, because waiting a few seconds will not help. Distinguish "slow down" from "you have hit a hard cap" in your client.

Rule 1: respect Retry-After

When a 429 includes Retry-After, wait at least that long before retrying that request. Do not retry immediately, and do not let every worker retry at the same instant when the wait ends.

import random, time, requests

def send_with_retry(session, payload, idem_key, max_attempts=6):
    delay = 1.0
    for attempt in range(1, max_attempts + 1):
        resp = session.post(
            "https://api.example.com/v1/emails",
            json=payload,
            headers={"Idempotency-Key": idem_key},
            timeout=10,
        )
        if resp.status_code != 429 and resp.status_code < 500:
            return resp
        retry_after = resp.headers.get("Retry-After")
        wait = float(retry_after) if retry_after and retry_after.isdigit() else delay
        time.sleep(wait + random.uniform(0, wait * 0.25))   # jitter
        delay = min(delay * 2, 60)
    raise RuntimeError("send failed after retries")

Note what this sketch does: it uses Retry-After when present, falls back to exponential backoff when not, adds jitter so workers spread out, caps the delay, and gives up after a bounded number of attempts. Production code would also handle the HTTP-date form of Retry-After and network exceptions.

Rule 2: keep retries idempotent

Rate limiting and retries go together, and retries are where duplicate emails come from. Consider a request that times out on your side after the provider accepted it. Your client sees an error and retries; without protection, the customer gets two messages.

If your provider supports an idempotency key header, send one with every request, generated once per logical message and reused on every retry of that message. Then a retry of a request that actually succeeded returns the original result instead of sending again. Never generate a fresh key inside the retry loop; that defeats the purpose.

Rule 3: smooth bursts before they reach the API

The best way to handle 429s is to rarely see them. Put a client-side rate limiter in front of the API, shared by all workers that send email.

A token bucket is the classic approach. The bucket holds up to N tokens and refills at R tokens per second. Each request takes a token; if none are available, the request waits. It allows short bursts up to N while holding the average rate to R.

For a single process, an in-memory limiter is enough. For many workers across machines, use a shared store (Redis is common) or route all sends through a queue consumed by a fixed number of workers. A queue with a known worker count and per-worker rate is often the simplest reliable design: total throughput is workers multiplied by per-worker rate, and it never exceeds what you configured.

Set your client-side rate somewhat below the provider's published limit to leave headroom for retries and other services sharing the key.

Rule 4: prioritize

When you are rate-limited, not every email is equally urgent. A password reset or a sign-in code matters now; a weekly digest can wait an hour. Use separate queues or priorities:

  • High priority: authentication codes, password resets, security alerts, payment receipts.
  • Normal: notifications triggered by user actions.
  • Low: digests, reminders, bulk announcements.

Workers drain high-priority queues first. During a backlog, critical mail keeps flowing while bulk mail waits.

Rule 5: watch for retry storms

A retry storm happens when many clients retry at once, keeping the API overloaded and producing more 429s and more retries. Prevent it with:

  • Jitter on every backoff.
  • A cap on total in-flight retries.
  • A circuit breaker that pauses sending for a short period after a burst of 429s or 5xx responses, rather than letting every worker hammer the API.

Monitoring

Track and alert on:

  • 429 responses per minute, by API key and endpoint.
  • Queue depth and age of the oldest message, by priority.
  • Time from enqueue to provider acceptance, especially for high-priority mail.
  • Requests abandoned after maximum retries.

A rising 429 rate is often the first sign that traffic patterns changed: a new feature sending more mail, or an import job nobody coordinated.

Checklist

  • Treat 429 as "not processed, retry later," and honor Retry-After.
  • Use exponential backoff with jitter and a cap when no header is provided.
  • Send a stable idempotency key per logical message on every attempt.
  • Rate-limit on the client side, shared across all workers, below the provider's limit.
  • Separate critical and bulk email into prioritized queues.
  • Distinguish temporary rate limits from hard quota errors.
  • Monitor 429 rates, queue age and abandoned sends.

Key takeaways

  • A 429 means nothing was sent; wait, then retry the same request.
  • Honor Retry-After, add jitter, and cap retries to avoid storms.
  • Idempotency keys prevent retries from becoming duplicate emails.
  • Client-side rate limiting and priority queues mean you rarely see a 429, and critical mail still flows when you do.

Start with Koltrix

Your domain, one inbox, and an API that sends.

A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

SharePost on XLinkedIn