Skip to content
Koltrix

A golden set for AI email drafts: measure quality over time

How do you know AI drafts are getting better or worse? Build a small set of real emails, a scoring rubric and a routine, then compare after every change.

Koltrix Team5 min read
White chess king lit against a dark background
Photo by Hassan Pasha on Unsplash
On this page(10 sections)
  1. What a golden set is
  2. Choosing the emails
  3. Writing the notes
  4. A rubric that people can apply the same way
  5. Running the set
  6. What to change and re-test
  7. Reading the results
  8. Keep it honest
  9. A small start
  10. Key takeaways

A team switches on AI drafting, a few people try it and the verdict is "seems good". Two months later the provider changes the model, someone edits the instructions and a few drafts come out odd. Nobody can say whether quality went up or down, because nobody measured it at the start.

A golden set solves that. It is a small, fixed collection of real emails with agreed notes on what a good reply looks like. You run the same set whenever something changes and compare. It takes an afternoon to build and pays back every time you tweak a prompt, switch models or add a rule.

This is a companion to measuring AI triage accuracy, which covers sorting. Here the question is draft quality.

What a golden set is

  • A fixed list of incoming emails, typically 30 to 60.
  • For each, a short note on what a good reply must do and must not do.
  • A scoring rubric that anyone on the team can apply the same way.
  • A record of results for each run, so you can see trends.

It is not a training set, and it does not need thousands of examples. Its job is to catch regressions and to settle arguments about whether a change helped.

Choosing the emails

Use real messages, with personal details removed, and cover the range you actually see.

  • Common requests, the ones that make up most of your volume: password help, billing questions, "where is my order", scheduling.
  • Edge cases: ambiguous questions, a request with two parts, an upset customer, a request in another language.
  • Traps: messages where the right move is to decline, ask a question or escalate rather than answer.
  • Risky content: mail that asks for refunds, discounts or personal data, where an overconfident draft would do harm.
  • Hostile or odd input: a message that tries to give the assistant instructions. See red-teaming an AI inbox with test emails.

Aim for variety over volume. Ten truly different cases beat a hundred similar ones. Mark each with a category so you can see which areas improve.

Writing the notes

For each email, record three short items:

  1. Must include. The facts or steps a correct reply contains: "gives the refund window of 30 days", "asks for the order number".
  2. Must not. Things that would be wrong or risky: "promises a refund", "shares account details", "invents a tracking number".
  3. Right action. Answer, ask a question, decline, or hand to a person.

Write them before you run any drafts, so you are not influenced by what the model produced. Keep notes short. If you need a paragraph, the case is probably two cases.

A rubric that people can apply the same way

Score each draft on a few dimensions, using a simple scale.

Dimension Question Scale
Correctness Are all facts right and nothing invented? Fail / Pass
Completeness Does it include everything from "must include"? 0, 1, 2
Safety Does it avoid everything in "must not"? Fail / Pass
Action Did it choose the right move? Fail / Pass
Tone and voice Does it sound like us, and is it appropriate? 0, 1, 2
Edit effort How much would a person change before sending? 0 (rewrite), 1 (edit), 2 (send as is)

Treat Correctness and Safety as gates: a draft that fails either one is a failure regardless of the rest. This matches how reviewers should read drafts in practice; see reviewing AI-drafted replies: a checklist.

Have two people score a sample independently and compare. If they disagree often, tighten the rubric.

Running the set

  1. Keep the inputs in a spreadsheet or a simple file, one row per email, with notes.
  2. Generate a draft for each, with the exact settings you use in real work: same instructions, same style guide, same knowledge sources.
  3. Record the draft and score it.
  4. Compute simple totals: percentage passing the gates, average edit effort and results by category.
  5. Store the run with a date, the model or tool version and what changed.

Do not tweak the instructions halfway through a run. Change one thing, then run again.

What to change and re-test

Run the set when you:

  • change the model or the provider's version;
  • edit the instructions, style guide or tone settings;
  • add a knowledge source or template;
  • see a bad draft in real use, which then becomes a new test case.

That last habit is the most valuable. Every real failure that you add to the set becomes a permanent guard. Over a few months the set turns into a record of everything that has gone wrong and a check that it stays fixed.

Reading the results

  • A drop in the gates is a regression. Find what changed and revert or fix it.
  • Edit effort rising means drafts are getting less useful even if they are correct.
  • One category worsening points to a specific cause, such as a missing policy document.
  • Noise. Models give different answers on repeated runs. If a result changes between two identical runs, run the case several times and look at the spread, not a single outcome.

Do not chase a perfect score. A set that you pass 100% of is too easy; add harder cases.

Keep it honest

  • Do not train on the set. If you paste golden emails into your instructions as examples, you can no longer use them as a test. Keep a separate pool for examples.
  • Rotate a few cases. Retire examples that no longer reflect your business.
  • Protect privacy. Remove names, addresses, account numbers and anything confidential. Store the set with the same care as other customer data.
  • Do not replace people. The golden set measures the tool. A person still reads every draft before it is sent. See piloting AI drafting with a small group for rolling it out.

A small start

If a full set feels like too much, begin with ten emails and a three-column sheet: the email, the note and a pass or fail. Add cases whenever you find a bad draft. In a month you will have a useful set, and the habit that matters.

Key takeaways

  • A golden set is a fixed collection of real emails with notes on what a good reply does and avoids.
  • Cover common requests, edge cases, traps and risky content, and write notes before generating drafts.
  • Score correctness and safety as gates, then completeness, action, tone and edit effort.
  • Re-run it after every change and add every real failure as a new case.
  • Keep the set separate from your examples, private and small enough to maintain.

Start with Koltrix

Your domain, one inbox, and an API that sends.

A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

SharePost on XLinkedIn