How AI email classification works, explained without hype
What an AI email classifier actually looks at, how rules and models split the work, why it gets things wrong, and how to judge one on your own inbox.

On this page(9 sections)
- What "classification" means for email
- The signals a classifier looks at
- 1. Headers you never see
- 2. The sender and your history with them
- 3. The content
- 4. Structure and formatting
- Rules and models: who does what
- Confidence, and why it matters
- Why classifiers get things wrong
- Fewer categories work better
- How to judge a classifier on your own mail
- What classification should never do on its own
- Key takeaways
An AI that sorts your email is not reading your mind. It is making a fast, educated guess about which pile a message belongs in, based on signals that are mostly visible to you too. Once you know what those signals are, the results stop feeling like magic and its mistakes start to make sense.
This post walks through how email classification works in practice: the inputs, how rules and models divide the job, where confidence comes in, and why the number of categories matters more than most people expect.
What "classification" means for email
Classification is a narrow task. Given one message (and sometimes its thread), pick one label from a fixed list. That's all. It is not summarizing, not drafting, and not deciding what to do. It answers a single question: which bucket?
The fixed list is the important part. A classifier choosing between five clear options behaves very differently from one choosing between forty overlapping ones. We'll come back to that.
The signals a classifier looks at
Most of the useful evidence is in the message itself. Roughly in order of how much it tends to help:
1. Headers you never see
Every email carries technical headers, and several of them are strong hints:
List-UnsubscribeandList-Idshow up on almost all bulk mail. Their presence alone is a decent hint that something is a newsletter or marketing send.Precedence: bulkorAuto-Submitted: auto-generatedmark automated messages.- Sending infrastructure. Messages sent through bulk email platforms carry recognizable patterns in
ReceivedandReturn-Pathheaders. - Authentication results. Whether SPF, DKIM and DMARC passed tells you whether the sender is who they claim to be, which matters for anything that looks like it came from a bank or a colleague.
Here is a trimmed example of what a typical product newsletter carries:
From: "Acme Weekly" <[email protected]>
List-Unsubscribe: <https://acme.example/u/abc123>, <mailto:[email protected]>
List-Unsubscribe-Post: List-Unsubscribe=One-Click
Precedence: bulk
A model doesn't need to read a word of the body to have a strong prior about this one.
2. The sender and your history with them
Has anyone on your team ever replied to this address? Is the sender's domain the same as a customer's? Is it a free webmail address writing for the first time? Past interaction is one of the strongest signals for "this is a real conversation", and its absence is a hint for cold outreach.
3. The content
This is where language models earn their keep. The body and subject carry intent: a question, a complaint, an invoice, a sales pitch, a password reset. Older systems used keyword lists ("unsubscribe", "invoice", "limited time"), which are brittle. Modern models read the text more like a person would, so "Quick question about your API limits" and "Hey, is there a cap on requests per minute?" land in the same place despite sharing no keywords.
4. Structure and formatting
Heavy HTML with images and tracking pixels, a single call-to-action button, a footer full of legal text and a postal address: all of this looks like marketing. A short plain-text note with a signature looks like a person.
Rules and models: who does what
Good classification systems use both, because they fail in different ways.
| Rules | Model | |
|---|---|---|
| Example | "Anything from [email protected] gets the Finance label" | "Is this a cold sales pitch?" |
| Strength | Predictable, explainable, instant | Handles wording it has never seen |
| Weakness | Breaks when senders change addresses or phrasing | Occasionally confident and wrong |
| Best for | Known senders, exact patterns | Fuzzy intent, first-time senders |
A sensible order is: deterministic rules first (they encode things you know), then the model for everything the rules didn't catch. When a rule and a model disagree, the rule should usually win, because someone wrote it on purpose.
Confidence, and why it matters
A classifier doesn't just output a label; internally it has some measure of how sure it is. A message with List-Unsubscribe, a marketing template and no reply history is an easy call. A short message from an unknown gmail address saying "following up on our chat" is genuinely ambiguous: it could be a customer, a recruiter or a cold pitch pretending to be a follow-up.
Well-designed systems treat low-confidence cases conservatively. If the classifier isn't sure, the safer error is to leave the message in the main view rather than bury it. Hiding a real customer email in a low-priority pile costs far more than letting one extra newsletter through.
Why classifiers get things wrong
Mistakes are not random. They cluster around a few predictable situations:
- People using personal addresses for business. A customer writing from their personal gmail looks a lot like a stranger.
- Legitimate senders using bulk tools. An investor update or a customer's announcement sent through a newsletter platform carries all the bulk-mail headers.
- Cold pitches disguised as replies. "Re:" in the subject and "as discussed" in the body are common tricks.
- Automated mail that needs action. A failed-payment notice or a security alert is automated, but it isn't noise.
- Forwarded threads. The visible sender is your colleague, but the content belongs to someone else.
None of these are reasons to avoid AI sorting. They are reasons to keep categories reviewable, make moving a message one click, and look at the low-priority piles now and then.
Fewer categories work better
It's tempting to ask for fine-grained buckets: Sales, Partnerships, Investors, Press, Recruiting, Vendors, Customers-Tier-1. Each extra category creates new borders, and every border is a place where the model (and frankly, a human) can disagree with itself.
A small set of mutually exclusive categories, each with an obvious meaning, gives you much more consistent results. Koltrix, for example, sorts incoming mail into five: Primary, Other, Cold pitches, Newsletters and Updates. Finer distinctions are better handled by labels and rules you control, layered on top.
The test for a good category is simple: could two people on your team look at the same email and agree on where it goes, quickly, most of the time? If not, the category is too vague for a model too.
How to judge a classifier on your own mail
Vendor demos use clean examples. Your inbox isn't clean. A quick, honest test:
- Pick a recent week of real incoming mail.
- Note where the classifier put each message and where you would have put it.
- Separate costly errors from cheap ones. A customer email filed as a newsletter is costly. A newsletter left in the main view is cheap.
- Check the edge cases listed above specifically: personal-address customers, automated alerts, disguised pitches.
- See how easy correction is. Can you move a message in one click? Can you add a rule so it never happens again for that sender?
If costly errors are rare and corrections are easy, the classifier is doing its job.
What classification should never do on its own
Sorting is low-risk because it's reversible: a misfiled email can be moved back. That's exactly why classification is a good first job for AI in an inbox. Actions that aren't reversible, like sending a reply, deserve a different standard. In Koltrix, AI sorts, summarizes and drafts, but nothing is sent without a person clicking send.
Key takeaways
- Email classifiers rely mostly on visible signals: headers like
List-Unsubscribe, sender history, content and formatting. - Rules handle what you know for certain; models handle fuzzy intent. Use both, rules first.
- Errors cluster in predictable places, so check those cases deliberately.
- A few clear, non-overlapping categories beat many fine-grained ones.
- Prefer conservative behavior: when unsure, keep a message visible rather than bury it.
Start with Koltrix
Your domain, one inbox, and an API that sends.
A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

