Per-provider throttling: pacing mail to big mailbox providers
Gmail, Outlook and Yahoo each defer differently when you send too fast. Build per-destination queues and adaptive concurrency that respond to 4xx replies.

On this page(10 sections)
Sending 50,000 messages as fast as your workers can push them is a good way to get 40,000 of them deferred. Large mailbox providers each decide how much mail they will accept from you, how fast, and over how many connections.
A sending system that respects those limits per destination delivers faster overall than one that ignores them.
Why one global send rate does not work
Your outbound mail goes to many receiving organizations, and each runs its own acceptance policies. Limits typically depend on:
- Your sending IP's and domain's reputation with that provider.
- How many simultaneous connections you open.
- How many messages you send per connection and per minute.
- How recipients at that provider have reacted to your recent mail.
None of the large providers publish fixed numbers, and their behavior changes with your reputation. What they do publish, in effect, is feedback: when you exceed what they will accept, they respond with temporary 4xx errors. A well-designed sender treats those responses as a control signal.
A single global rate fails in both directions. Set it high, and you flood one provider while others are fine. Set it low enough for the strictest provider, and every other destination waits for no reason.
Group destinations by provider, not by domain
Recipient domains are not the unit that matters. Thousands of custom domains route their mail to the same large providers. The useful grouping is the infrastructure that receives the mail, which you can determine from MX records:
example-customer.com. MX 10 mx.mailhost-a.example.
another-company.com. MX 10 mx.mailhost-a.example.
Both domains deliver to the same receiving infrastructure, so they should share a throttle. In practice, teams maintain a mapping from MX hostname patterns to provider groups, with a default group for everything else:
| Provider group | Matched by | Starting concurrency | Notes |
|---|---|---|---|
| Large provider A | Known MX host patterns | Moderate | Strong reputation needed for higher rates |
| Large provider B | Known MX host patterns | Conservative | Sensitive to bursts |
| Large provider C | Known MX host patterns | Moderate | |
| Default | Everything else | Low per destination | Many small receivers |
The starting values are yours to tune. They should be conservative for new IPs and grow as results allow.
Per-group queues
The core design is one logical queue per provider group, each with its own limits:
- Max concurrent connections to that provider.
- Max messages per connection before reconnecting.
- Max messages per minute as a rate limit.
- Backoff state that tightens limits after deferrals and relaxes them after successes.
Workers pull from whichever queues have available capacity. A slowdown at one provider fills only that provider's queue; everything else keeps flowing.
Adaptive concurrency
Static limits are a starting point. A better system adjusts them using the provider's responses, similar to how TCP congestion control works:
class ProviderThrottle:
def __init__(self, start=4, minimum=1, maximum=40):
self.limit = start
self.minimum, self.maximum = minimum, maximum
self.successes = 0
def on_accepted(self):
self.successes += 1
if self.successes >= self.limit * 20: # sustained success
self.limit = min(self.maximum, self.limit + 1)
self.successes = 0
def on_deferred(self, code: str, text: str):
self.successes = 0
if code.startswith("421") or "rate" in text.lower():
self.limit = max(self.minimum, self.limit // 2) # back off sharply
Additive increase, multiplicative decrease: grow slowly while things work, cut quickly when the provider pushes back. Add a cooldown so a single deferral does not oscillate the limit up and down.
Reading the provider's signals
Not every 4xx means the same thing:
- Rate or connection limits ("too many connections," "try again later," "rate limited"). Reduce concurrency and messages per minute for that provider.
- Reputation-related deferrals (text mentioning unusual volume, user complaints or policy). Reduce volume more aggressively and check what changed in your sending.
- Greylisting (first-attempt deferrals from smaller servers). Retry after the normal delay; no throttle change needed.
- Mailbox-specific (mailbox full, temporarily unavailable). Retry that recipient later; do not throttle the whole provider.
Exact wording differs by provider and changes over time, so match broadly and log the raw response text. Enhanced status codes (RFC 3463), such as 4.7.x for policy-related issues, help with classification where providers include them.
Prioritize within each queue
When a provider's capacity is limited, decide which mail goes first. A verification code waiting behind a newsletter is a product bug. Use priority levels inside each provider queue:
- Security mail: verification codes, password resets, sign-in links.
- Transactional mail: receipts, alerts.
- Notifications and digests.
- Marketing.
Better still, send marketing from separate infrastructure or IPs so its throttling never competes with critical mail.
Retry timing
Throttling decides how fast you send. Retry policy decides when to try again after a deferral. Use exponential backoff with jitter per message, and a give-up window long enough to ride out provider incidents. RFC 5321 suggests that senders keep retrying for at least four to five days before giving up, though for time-sensitive mail like verification codes a much shorter useful lifetime often makes more sense; an hour-old code is rarely worth delivering.
What to monitor
- Deferral rate per provider group.
- Current concurrency limit per provider, graphed over time.
- Queue depth and oldest-message age per provider.
- Delivery latency per priority level.
A provider whose concurrency limit keeps getting cut is telling you something about your reputation there, not just about capacity.
Checklist
- Group destinations by receiving infrastructure (MX), not by recipient domain.
- Maintain a queue per provider group with its own concurrency and rate limits.
- Adapt limits: increase slowly on success, halve on rate-limit deferrals.
- Classify 4xx responses before reacting; log raw text.
- Prioritize security and transactional mail within each queue.
- Monitor per-provider deferrals, limits, queue age and latency.
Key takeaways
- Each mailbox provider sets its own acceptance limits, and they depend on your reputation.
- Per-provider queues keep a slowdown at one receiver from delaying everyone else.
- Adaptive concurrency, grow slowly and cut quickly, uses deferrals as a control signal.
- Classify deferrals: rate limits, reputation, greylisting and mailbox issues need different responses.
- Put critical mail first, ideally on infrastructure separate from marketing.
Start with Koltrix
Your domain, one inbox, and an API that sends.
A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.


