Measuring AI triage accuracy in your own inbox
A worked example of testing AI email sorting on your own mail: sample 200 messages, label them by hand, build a confusion matrix and weigh costly errors.

On this page(10 sections)
- Step 1: Decide what you are measuring
- Step 2: Pull a representative sample
- Step 3: Label by hand, blind if you can
- Step 4: Build the confusion matrix
- Step 5: Compute precision and recall per category
- Step 6: Weight errors by what they cost
- Step 7: Turn findings into fixes
- Step 8: Repeat on a schedule
- A checklist for your own run
- Bottom line
"It feels pretty accurate" is not a measurement. If your team is going to trust an AI to decide what lands in front of them each morning, spend one afternoon finding out how often it is right on your mail, and which mistakes it makes.
This is a worked example. We'll follow a hypothetical five-person SaaS team through sampling, hand-labeling, scoring and interpreting the results. You need a spreadsheet, a couple of hours, and the willingness to be honest about edge cases.
Step 1: Decide what you are measuring
Before touching any email, write down the categories the AI chooses between and a one-line definition for each. Our example team uses five:
| Category | Definition used for hand-labeling |
|---|---|
| Primary | A person (or system) expecting a response or action from us |
| Other | Real but low priority: FYIs, receipts for things we expected |
| Cold pitches | Unsolicited sales or partnership outreach |
| Newsletters | Bulk content we subscribed to |
| Updates | Automated notifications from tools and services |
The definitions matter more than they look. If two humans can't agree on what "Other" means, the AI can't be scored fairly against either of them.
Step 2: Pull a representative sample
Take roughly 200 incoming messages. Two hundred is enough to see patterns and small enough to label in an hour or so. A few rules for the sample:
- Use a continuous window, such as the last full week, rather than picking interesting messages. Cherry-picking makes any system look worse or better than it is.
- Include weekends if mail arrives then; the mix often differs.
- Exclude mail you have already corrected, or note it separately, because your corrections may have influenced later sorting.
- Cover every shared mailbox the AI sorts, not just your personal one.
Export or copy into a sheet: sender, subject, date, and the category the AI assigned. Leave a blank column for your label.
Step 3: Label by hand, blind if you can
Hide the AI's column while you label. Seeing its answer first nudges you toward agreeing with it. If two people can label the same 200, even better: where you disagree with each other, the email is genuinely ambiguous, and you shouldn't count the AI's choice there as a hard error.
Our example team had two people label independently. They disagreed on 11 messages, mostly vendor emails that were half update, half pitch. They talked those through, settled on a label, and flagged them as "ambiguous" for later.
Step 4: Build the confusion matrix
A confusion matrix is just a grid: rows are what the message really was (your label), columns are what the AI said. The diagonal is where they agree.
Here is our example team's result for 200 messages (illustrative numbers):
| Actual \ AI said | Primary | Other | Cold | News | Updates | Total |
|---|---|---|---|---|---|---|
| Primary | 41 | 3 | 1 | 1 | 2 | 48 |
| Other | 4 | 17 | 0 | 2 | 3 | 26 |
| Cold | 3 | 1 | 22 | 0 | 0 | 26 |
| News | 0 | 1 | 1 | 44 | 2 | 48 |
| Updates | 1 | 2 | 0 | 3 | 46 | 52 |
| Total | 49 | 24 | 24 | 50 | 53 | 200 |
The diagonal sums to 170 of 200. That single "overall accuracy" figure is the least useful number in the table, because it treats every mistake as equal.
Step 5: Compute precision and recall per category
For each category, two questions matter:
- Recall: of the messages that truly belong here, how many did the AI catch? (diagonal cell ÷ row total)
- Precision: of the messages the AI put here, how many belong? (diagonal cell ÷ column total)
For the example team:
Category Recall Precision
Primary 41/48 = 0.85 41/49 = 0.84
Other 17/26 = 0.65 17/24 = 0.71
Cold pitches 22/26 = 0.85 22/24 = 0.92
Newsletters 44/48 = 0.92 44/50 = 0.88
Updates 46/52 = 0.88 46/53 = 0.87
"Other" is the weakest category on both measures. That's common: a catch-all bucket has the fuzziest definition, so both humans and models disagree about it.
Step 6: Weight errors by what they cost
Now the important part. Look at the off-diagonal cells and ask what each mistake would have cost you.
Costly errors are real conversations that got hidden. In our table:
- 1 Primary message filed as Cold pitch
- 1 Primary message filed as Newsletter
- 2 Primary messages filed as Updates
- 3 Primary messages filed as Other
That's 7 of 48 important messages not shown first. Open each one. In the example, two were customers writing from personal addresses, one was a billing alert from the payment processor, and one was an investor update sent through a newsletter tool. The billing alert was the worst miss: it needed action within days.
Cheap errors are noise that leaked into Primary: 4 Other, 3 Cold and 1 Update shown as Primary. Annoying, but you see them and move on in seconds.
A system with slightly lower overall accuracy but fewer costly errors is the better system for most teams. Hiding a customer is worse than showing a pitch.
Step 7: Turn findings into fixes
Each costly error should lead to one of three actions:
- A deterministic rule when the pattern is exact. The payment processor's alert address can always go to Primary. Rules beat models for known senders.
- A definition change when humans disagreed too. If "Other" keeps confusing everyone, tighten its definition or accept that it's a soft bucket and skim it daily.
- A habit when no rule can fix it. Customers on personal addresses will always be hard to spot; a quick daily glance at Other and Updates catches them.
Record what you changed so the next measurement can show whether it helped.
Step 8: Repeat on a schedule
Mail mix drifts. New tools start sending notifications, a newsletter changes platforms, your customer base shifts. Re-run a smaller version (50 to 100 messages) every quarter, or after any big change in what you receive. Keep the spreadsheets; trends across measurements tell you more than any single snapshot.
A checklist for your own run
- Categories written down with one-line definitions
- Continuous sample of about 200 messages across all sorted mailboxes
- Hand labels made without seeing the AI's answer
- Ambiguous messages flagged and set aside
- Confusion matrix built
- Recall and precision per category
- Every important-message miss opened and explained
- One fix (rule, definition or habit) per costly error pattern
- Date set for the next measurement
Bottom line
Measure on your own mail, not a demo inbox. Build the confusion matrix, but judge the results by which errors hide real conversations rather than by one accuracy figure. Most costly misses fall into a few repeatable patterns, and each pattern usually has a simple fix: a rule, a clearer definition, or a short daily habit.
Start with Koltrix
Your domain, one inbox, and an API that sends.
A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.


