Skip to content

Measuring AI triage accuracy in your own inbox

A worked example of testing AI email sorting on your own mail: sample 200 messages, label them by hand, build a confusion matrix and weigh costly errors.

Koltrix Team5 min read
Performance analytics graphs on a laptop screen
Photo by Luke Chesser on Unsplash
On this page(10 sections)
  1. Step 1: Decide what you are measuring
  2. Step 2: Pull a representative sample
  3. Step 3: Label by hand, blind if you can
  4. Step 4: Build the confusion matrix
  5. Step 5: Compute precision and recall per category
  6. Step 6: Weight errors by what they cost
  7. Step 7: Turn findings into fixes
  8. Step 8: Repeat on a schedule
  9. A checklist for your own run
  10. Bottom line

"It feels pretty accurate" is not a measurement. If your team is going to trust an AI to decide what lands in front of them each morning, spend one afternoon finding out how often it is right on your mail, and which mistakes it makes.

This is a worked example. We'll follow a hypothetical five-person SaaS team through sampling, hand-labeling, scoring and interpreting the results. You need a spreadsheet, a couple of hours, and the willingness to be honest about edge cases.

Step 1: Decide what you are measuring

Before touching any email, write down the categories the AI chooses between and a one-line definition for each. Our example team uses five:

Category Definition used for hand-labeling
Primary A person (or system) expecting a response or action from us
Other Real but low priority: FYIs, receipts for things we expected
Cold pitches Unsolicited sales or partnership outreach
Newsletters Bulk content we subscribed to
Updates Automated notifications from tools and services

The definitions matter more than they look. If two humans can't agree on what "Other" means, the AI can't be scored fairly against either of them.

Step 2: Pull a representative sample

Take roughly 200 incoming messages. Two hundred is enough to see patterns and small enough to label in an hour or so. A few rules for the sample:

  • Use a continuous window, such as the last full week, rather than picking interesting messages. Cherry-picking makes any system look worse or better than it is.
  • Include weekends if mail arrives then; the mix often differs.
  • Exclude mail you have already corrected, or note it separately, because your corrections may have influenced later sorting.
  • Cover every shared mailbox the AI sorts, not just your personal one.

Export or copy into a sheet: sender, subject, date, and the category the AI assigned. Leave a blank column for your label.

Step 3: Label by hand, blind if you can

Hide the AI's column while you label. Seeing its answer first nudges you toward agreeing with it. If two people can label the same 200, even better: where you disagree with each other, the email is genuinely ambiguous, and you shouldn't count the AI's choice there as a hard error.

Our example team had two people label independently. They disagreed on 11 messages, mostly vendor emails that were half update, half pitch. They talked those through, settled on a label, and flagged them as "ambiguous" for later.

Step 4: Build the confusion matrix

A confusion matrix is just a grid: rows are what the message really was (your label), columns are what the AI said. The diagonal is where they agree.

Here is our example team's result for 200 messages (illustrative numbers):

Actual \ AI said Primary Other Cold News Updates Total
Primary 41 3 1 1 2 48
Other 4 17 0 2 3 26
Cold 3 1 22 0 0 26
News 0 1 1 44 2 48
Updates 1 2 0 3 46 52
Total 49 24 24 50 53 200

The diagonal sums to 170 of 200. That single "overall accuracy" figure is the least useful number in the table, because it treats every mistake as equal.

Step 5: Compute precision and recall per category

For each category, two questions matter:

  • Recall: of the messages that truly belong here, how many did the AI catch? (diagonal cell ÷ row total)
  • Precision: of the messages the AI put here, how many belong? (diagonal cell ÷ column total)

For the example team:

Category      Recall          Precision
Primary       41/48 = 0.85    41/49 = 0.84
Other         17/26 = 0.65    17/24 = 0.71
Cold pitches  22/26 = 0.85    22/24 = 0.92
Newsletters   44/48 = 0.92    44/50 = 0.88
Updates       46/52 = 0.88    46/53 = 0.87

"Other" is the weakest category on both measures. That's common: a catch-all bucket has the fuzziest definition, so both humans and models disagree about it.

Step 6: Weight errors by what they cost

Now the important part. Look at the off-diagonal cells and ask what each mistake would have cost you.

Costly errors are real conversations that got hidden. In our table:

  • 1 Primary message filed as Cold pitch
  • 1 Primary message filed as Newsletter
  • 2 Primary messages filed as Updates
  • 3 Primary messages filed as Other

That's 7 of 48 important messages not shown first. Open each one. In the example, two were customers writing from personal addresses, one was a billing alert from the payment processor, and one was an investor update sent through a newsletter tool. The billing alert was the worst miss: it needed action within days.

Cheap errors are noise that leaked into Primary: 4 Other, 3 Cold and 1 Update shown as Primary. Annoying, but you see them and move on in seconds.

A system with slightly lower overall accuracy but fewer costly errors is the better system for most teams. Hiding a customer is worse than showing a pitch.

Step 7: Turn findings into fixes

Each costly error should lead to one of three actions:

  1. A deterministic rule when the pattern is exact. The payment processor's alert address can always go to Primary. Rules beat models for known senders.
  2. A definition change when humans disagreed too. If "Other" keeps confusing everyone, tighten its definition or accept that it's a soft bucket and skim it daily.
  3. A habit when no rule can fix it. Customers on personal addresses will always be hard to spot; a quick daily glance at Other and Updates catches them.

Record what you changed so the next measurement can show whether it helped.

Step 8: Repeat on a schedule

Mail mix drifts. New tools start sending notifications, a newsletter changes platforms, your customer base shifts. Re-run a smaller version (50 to 100 messages) every quarter, or after any big change in what you receive. Keep the spreadsheets; trends across measurements tell you more than any single snapshot.

A checklist for your own run

  • Categories written down with one-line definitions
  • Continuous sample of about 200 messages across all sorted mailboxes
  • Hand labels made without seeing the AI's answer
  • Ambiguous messages flagged and set aside
  • Confusion matrix built
  • Recall and precision per category
  • Every important-message miss opened and explained
  • One fix (rule, definition or habit) per costly error pattern
  • Date set for the next measurement

Bottom line

Measure on your own mail, not a demo inbox. Build the confusion matrix, but judge the results by which errors hide real conversations rather than by one accuracy figure. Most costly misses fall into a few repeatable patterns, and each pattern usually has a simple fix: a rule, a clearer definition, or a short daily habit.

Start with Koltrix

Your domain, one inbox, and an API that sends.

A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

SharePost on XLinkedIn