Skip to content

Unicode in email: SMTPUTF8, encoded headers and IDNs

Non-ASCII names, subjects and addresses need different mechanisms. How RFC 2047 encoded words, SMTPUTF8 and IDN domains fit together in practice.

Koltrix Team4 min read
A teal LED panel with abstract light patterns
Photo by Adi Goldstein on Unsplash
On this page(8 sections)
  1. Four places non-ASCII text can appear
  2. Bodies: charset and transfer encoding
  3. Headers: RFC 2047 encoded words
  4. Domains: internationalized domain names
  5. Local parts: SMTPUTF8
  6. What this means for a SaaS application
  7. Accepting addresses at signup
  8. Sending
  9. Testing
  10. Checklist
  11. Key takeaways

Email was designed for seven-bit ASCII, and the world writes in thousands of scripts. Over the years, three separate mechanisms were bolted on to handle non-ASCII text in different parts of a message.

Knowing which one applies where explains most internationalization bugs in transactional email.

Four places non-ASCII text can appear

Location Example Mechanism
Message body "Ihre Bestellung ist unterwegs" with umlauts MIME charset plus transfer encoding
Header text (Subject, display names) Subject: Votre reçu RFC 2047 encoded words
Domain part of an address user@bücher.example IDNA (internationalized domain names)
Local part of an address josé@example.com SMTPUTF8 (RFC 6531)

Each mechanism solves a different problem, and they have very different levels of support.

Bodies: charset and transfer encoding

This is the solved problem. A MIME part declares its character set, and a transfer encoding makes the bytes safe for transport:

Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

Ihre Bestellung ist unterwegs. Vielen Dank f=C3=BCr Ihren Einkauf.

Use UTF-8 for everything. Choose quoted-printable for text that is mostly ASCII (readable in raw form) or base64 for text that is mostly non-ASCII (more compact). Servers that advertise 8BITMIME can accept raw 8-bit bodies, but pre-encoding avoids relays converting your message in transit, which would also break DKIM signatures.

The classic bug here is a mismatch: UTF-8 bytes declared as ISO-8859-1, producing mojibake such as für instead of für. Always declare the charset you actually used.

Headers: RFC 2047 encoded words

Header fields were historically ASCII-only. RFC 2047 introduced "encoded words" so that Subject lines and display names can carry other characters:

Subject: =?UTF-8?Q?Votre_re=C3=A7u_n=C2=B0_A-10293?=
From: =?UTF-8?B?Sm9zw6kgR2FyY8OtYQ==?= <[email protected]>

The format is =?charset?encoding?text?=, where encoding is Q (a quoted-printable variant in which underscores represent spaces) or B (base64). Each encoded word is limited to 75 characters, so long subjects are split into several encoded words, separated by folding whitespace.

You should almost never write these by hand. Every mainstream mail library encodes headers automatically when given Unicode strings. Bugs appear when code assembles headers manually, splits an encoded word in the middle of a multibyte character, or double-encodes an already encoded value.

Important limitation: encoded words are only allowed in certain places, such as unstructured text (Subject) and display names. They may not be used inside the address itself. [email protected] is not a valid way to write a non-ASCII mailbox name.

Domains: internationalized domain names

Domain names with non-ASCII characters are handled by IDNA, which maps each Unicode label to an ASCII-compatible form starting with xn-- (called Punycode):

bücher.example  →  xn--bcher-kva.example

DNS itself only ever sees the ASCII form. For email, this means an address with an internationalized domain can be sent through ordinary SMTP by converting the domain part to its xn-- form, as long as the local part is ASCII:

RCPT TO:<[email protected]>

In your application, normalize and convert domain names with a maintained IDNA library rather than writing your own conversion. The current standard (IDNA2008) and the older IDNA2003 differ in some edge cases, and libraries differ in which they implement, so test the domains your customers actually use.

Store addresses consistently. Many systems store the Unicode form for display and compute the ASCII form at send time. Whatever you choose, compare addresses in a normalized form, or you will treat the same address as two different users.

Local parts: SMTPUTF8

The local part (before the @) is the hardest case. There is no ASCII encoding for it. RFC 6531 defines the SMTPUTF8 extension, which allows UTF-8 in envelope addresses and headers when every server on the path supports it.

The sending server checks whether the receiving server advertises the extension in its EHLO response, and if so adds a parameter to the transaction:

EHLO mail.example.com
250-mx.receiver.example
250-8BITMIME
250 SMTPUTF8
MAIL FROM:<[email protected]> SMTPUTF8
RCPT TO:<josé@example.net>

If the receiving server does not advertise SMTPUTF8, the message cannot be delivered to a non-ASCII local part. There is no downgrade path; it fails.

Support has grown, and several major mailbox providers accept SMTPUTF8, but it is far from universal across corporate gateways, older servers, sending providers and the application libraries in between. Every hop must support it, including your own provider's API and SMTP relay.

What this means for a SaaS application

Accepting addresses at signup

  • Accept internationalized domains and convert them correctly.
  • Decide deliberately whether to accept non-ASCII local parts. If your sending provider or downstream systems cannot deliver to them, accepting them creates accounts that never receive email. If you do accept them, test delivery end to end.
  • Apply Unicode normalization (NFC is a common choice) before storing and comparing.
  • Be alert to look-alike characters. Mixed-script addresses and domains can be used for impersonation; that is a concern for display and for security review, not a reason to reject all non-ASCII input.

Sending

  • Let your mail library encode Subject and display names.
  • Use UTF-8 and an explicit transfer encoding for bodies.
  • Convert domains to their ASCII form before handing addresses to SMTP if your library does not do it for you.
  • Check whether your provider supports SMTPUTF8 before promising delivery to non-ASCII local parts.

Testing

Include realistic international data in fixtures: names in several scripts, right-to-left text, emoji in subjects, an internationalized domain, and (if supported) a non-ASCII local part. Render and parse the messages in tests to catch double encoding and charset mismatches.

Checklist

  • Bodies are UTF-8 with quoted-printable or base64 transfer encoding.
  • Header text is encoded by a library, never assembled by hand.
  • Domains are converted with a maintained IDNA library and compared in normalized form.
  • Non-ASCII local parts are either supported end to end or rejected clearly at input.
  • Fixtures include multiple scripts, emoji and internationalized domains.

Key takeaways

  • Non-ASCII text in bodies, headers, domains and local parts uses four different mechanisms.
  • Bodies and header text are well supported; let libraries handle the encoding.
  • Internationalized domains travel as xn-- ASCII labels and work over ordinary SMTP.
  • Non-ASCII local parts need SMTPUTF8 on every hop, with no fallback, so verify support before accepting them.

Start with Koltrix

Your domain, one inbox, and an API that sends.

A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

SharePost on XLinkedIn