Skip to content

Outage communication over email: before, during, after

Copy-ready email templates for outages: first notice, updates, resolution and post-incident summary, plus cadence, who sends, and handling the inbound flood.

Koltrix Team5 min read
Rack of server equipment in a dark room
Photo by Tyler on Unsplash
On this page(12 sections)
  1. Decide these before you need them
  2. Internal coordination comes first
  3. Choosing who gets proactive email
  4. Stage 1: Initial notice
  5. Stage 2: Updates
  6. Stage 3: Resolution
  7. Stage 4: Post-incident summary
  8. Handling the inbound flood
  9. After the incident: the support side
  10. Cadence cheat sheet
  11. Mistakes to avoid
  12. Bottom line

During an outage, engineering fights the fire and support fights the inbox. The second fight is much easier if the emails were written before the fire started.

This post gives you templates for each stage of an incident, plus the decisions to make in advance: who sends, how often, and how to answer the flood of "is it down?" emails without writing each reply by hand.

Decide these before you need them

Write the answers into your support playbook now, when nobody is stressed.

  • Who declares an incident? Usually the engineering on-call. Support should know how to find out within minutes.
  • Who writes customer communication? One person, typically a support lead or the founder at a small company. Not the engineer fixing the problem.
  • Where is the source of truth? A status page if you have one. Email updates should match it word for word on facts.
  • Who gets proactive email? Everyone, only affected customers, or only accounts above a certain size? Proactive email to everyone for a ten-minute blip creates more alarm than it prevents.
  • What's the update cadence? A common default is every 30 to 60 minutes during an active incident, even if the update is "still working on it."

Internal coordination comes first

Customer emails are only as good as the information behind them. Set up one place where engineering posts what support is allowed to say, such as a dedicated incident channel with a pinned "current customer-facing status" message. The person writing customer updates reads from there and nowhere else.

This avoids two classic problems. The first is support repeating a theory an engineer mentioned in passing that later proves wrong. The second is support having nothing to say because engineering is too busy to answer questions. A single pinned status line, updated whenever it changes, solves both: engineers update one line, and support turns it into customer language.

Agree on vocabulary too. Words like "degraded," "partial outage" and "resolved" should mean the same thing to everyone on the team, and ideally match the labels on your status page.

Choosing who gets proactive email

Proactive outage email is a judgment call. Some guidelines that work for many small SaaS teams:

  • Short, contained incidents that customers might not even notice usually don't need proactive email. The status page and replies to people who write in are enough.
  • Longer incidents, or anything affecting data (lost messages, failed payments, missed scheduled jobs) usually do. Customers deserve to hear it from you rather than discover it later.
  • Partial incidents affecting one region, one plan, or one feature should go only to affected accounts when you can identify them. Emailing everyone about a problem most of them never saw creates needless worry.

Whatever you decide, send the resolution email to everyone who received the initial notice. Starting a story and never finishing it is worse than not starting it.

Stage 1: Initial notice

Send this once you've confirmed customer impact. Speed matters more than detail.

Subject: [Product] — service disruption affecting [feature]

Hi [first name or "there"],

We're currently seeing problems with [feature/area]. You may notice
[symptom customers will see: errors when loading dashboards, delayed
emails, failed logins].

Our team is investigating now. We'll send an update by [time + time zone],
or sooner if anything changes.

Live updates: [status page link]

[Name], [Company]

Notes:

  • Describe symptoms, not causes. "Dashboards may fail to load" is useful. "Database replication lag" isn't, and it may turn out to be wrong.
  • Commit to the next update time, not a fix time.
  • Don't speculate about duration. "Should be fixed within the hour" is the sentence most likely to be quoted back at you.

Stage 2: Updates

Send at the cadence you set, whether or not there's news.

Subject: Update: [Product] service disruption — [time]

Hi [first name or "there"],

Update on the [feature] problem:

- Current status: [investigating / cause identified / fix in progress /
  monitoring]
- What you may see: [symptoms, updated]
- What we're doing: [one or two plain sentences]
- Workaround, if any: [steps, or "none yet"]

Next update by [time + time zone].

[Name]

Keep updates short. The structure helps people scan, and it stops you from writing paragraphs of reassurance with no information in them. If nothing changed, say so plainly: "No change since our last update. The team is still working on a fix."

Stage 3: Resolution

Subject: Resolved: [Product] service disruption

Hi [first name or "there"],

The problem affecting [feature] has been resolved as of [time + time zone].
Everything should now be working normally.

[If customers need to act:] If you still see [symptom], please
[refresh / log out and back in / retry the failed action].

[If data was affected:] [What happened to queued items, e.g. "Emails queued
during the disruption have now been delivered."]

We'll share a short summary of what happened and what we're changing
within [timeframe].

Sorry for the disruption, and thank you for bearing with us.

[Name]

The resolution email is where a single, sincere apology belongs. Earlier updates should focus on information.

Stage 4: Post-incident summary

Send to the same audience, usually a few days later, once you understand the cause.

Subject: What happened during the [date] disruption

Hi [first name or "there"],

On [date], between [start] and [end] [time zone], [feature] was
[unavailable / degraded]. Here's a short account of what happened.

What customers experienced: [symptoms, duration]
What caused it: [plain-language explanation, two or three sentences]
How we fixed it: [one or two sentences]
What we're changing: [specific actions, e.g. added monitoring, changed
deployment process]

If this affected you in a way we haven't covered, reply to this email
and we'll look into it.

[Name], [Company]

Be specific about changes. "We're taking steps to prevent this" says nothing. "We now run a database migration check before every deploy" says something.

Handling the inbound flood

During an outage, many customers email at once with the same question. Prepare a canned status reply:

Hi [first name],

Thanks for letting us know. We're aware of a problem affecting [feature]
and are working on it now. You can follow live updates here: [status link].

We'll also email you when it's resolved. You don't need to do anything
in the meantime. [Workaround if any.]

[Name]

Tips for the flood:

  1. Label every incident-related email (for example incident-2026-09-17) so you can follow up with each person at resolution.
  2. Reply with the canned message, then move on. Save detailed troubleshooting for after the incident, when you can tell who is still affected.
  3. Watch for emails that are not about the outage. A billing question can hide among a hundred "is it down?" emails. Scan subjects for outliers.
  4. Reply to everyone labeled when it's resolved, even if they also got the proactive resolution email. They wrote to you; answer them.

After the incident: the support side

Once things are stable, a few support tasks remain that are easy to forget:

  • Follow up with anyone still affected. Some customers will have edge-case problems that outlast the main fix. The incident label makes them easy to find.
  • Check for billing consequences. If customers lost meaningful service, decide as a team whether credits are appropriate, and apply the decision consistently rather than only to whoever complains loudest.
  • Update your snippets and templates with anything you wished you'd had ready.
  • Contribute to the post-incident review. Support sees the customer impact engineering may not: which symptoms confused people, what questions came up repeatedly, which wording caused alarm.

Cadence cheat sheet

Stage When to send Length
Initial notice As soon as impact is confirmed Very short
Updates Every 30–60 minutes while active Short, structured
Resolution When fixed and verified Short
Post-incident summary Within a few days Medium

Mistakes to avoid

  • Silence between updates. No news is still news; send it.
  • Different facts in different places. Status page, email and social posts should agree.
  • Over-promising ETAs. Commit to update times instead.
  • Jargon. Customers need symptoms and actions, not internals.
  • Burying the apology in the first email. Lead with information; apologize at resolution.

Bottom line

Outage email is a writing problem you can solve in advance. Decide who sends and how often, keep the four templates ready, describe symptoms instead of guessing causes, promise update times rather than fix times, and label every inbound message so everyone hears back when it's over.

Start with Koltrix

Your domain, one inbox, and an API that sends.

A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

SharePost on XLinkedIn