A golden set for AI email drafts: measure quality over time
How do you know AI drafts are getting better or worse? Build a small set of real emails, a scoring rubric and a routine, then compare after every change.

On this page(10 sections)
A team switches on AI drafting, a few people try it and the verdict is "seems good". Two months later the provider changes the model, someone edits the instructions and a few drafts come out odd. Nobody can say whether quality went up or down, because nobody measured it at the start.
A golden set solves that. It is a small, fixed collection of real emails with agreed notes on what a good reply looks like. You run the same set whenever something changes and compare. It takes an afternoon to build and pays back every time you tweak a prompt, switch models or add a rule.
This is a companion to measuring AI triage accuracy, which covers sorting. Here the question is draft quality.
What a golden set is
- A fixed list of incoming emails, typically 30 to 60.
- For each, a short note on what a good reply must do and must not do.
- A scoring rubric that anyone on the team can apply the same way.
- A record of results for each run, so you can see trends.
It is not a training set, and it does not need thousands of examples. Its job is to catch regressions and to settle arguments about whether a change helped.
Choosing the emails
Use real messages, with personal details removed, and cover the range you actually see.
- Common requests, the ones that make up most of your volume: password help, billing questions, "where is my order", scheduling.
- Edge cases: ambiguous questions, a request with two parts, an upset customer, a request in another language.
- Traps: messages where the right move is to decline, ask a question or escalate rather than answer.
- Risky content: mail that asks for refunds, discounts or personal data, where an overconfident draft would do harm.
- Hostile or odd input: a message that tries to give the assistant instructions. See red-teaming an AI inbox with test emails.
Aim for variety over volume. Ten truly different cases beat a hundred similar ones. Mark each with a category so you can see which areas improve.
Writing the notes
For each email, record three short items:
- Must include. The facts or steps a correct reply contains: "gives the refund window of 30 days", "asks for the order number".
- Must not. Things that would be wrong or risky: "promises a refund", "shares account details", "invents a tracking number".
- Right action. Answer, ask a question, decline, or hand to a person.
Write them before you run any drafts, so you are not influenced by what the model produced. Keep notes short. If you need a paragraph, the case is probably two cases.
A rubric that people can apply the same way
Score each draft on a few dimensions, using a simple scale.
| Dimension | Question | Scale |
|---|---|---|
| Correctness | Are all facts right and nothing invented? | Fail / Pass |
| Completeness | Does it include everything from "must include"? | 0, 1, 2 |
| Safety | Does it avoid everything in "must not"? | Fail / Pass |
| Action | Did it choose the right move? | Fail / Pass |
| Tone and voice | Does it sound like us, and is it appropriate? | 0, 1, 2 |
| Edit effort | How much would a person change before sending? | 0 (rewrite), 1 (edit), 2 (send as is) |
Treat Correctness and Safety as gates: a draft that fails either one is a failure regardless of the rest. This matches how reviewers should read drafts in practice; see reviewing AI-drafted replies: a checklist.
Have two people score a sample independently and compare. If they disagree often, tighten the rubric.
Running the set
- Keep the inputs in a spreadsheet or a simple file, one row per email, with notes.
- Generate a draft for each, with the exact settings you use in real work: same instructions, same style guide, same knowledge sources.
- Record the draft and score it.
- Compute simple totals: percentage passing the gates, average edit effort and results by category.
- Store the run with a date, the model or tool version and what changed.
Do not tweak the instructions halfway through a run. Change one thing, then run again.
What to change and re-test
Run the set when you:
- change the model or the provider's version;
- edit the instructions, style guide or tone settings;
- add a knowledge source or template;
- see a bad draft in real use, which then becomes a new test case.
That last habit is the most valuable. Every real failure that you add to the set becomes a permanent guard. Over a few months the set turns into a record of everything that has gone wrong and a check that it stays fixed.
Reading the results
- A drop in the gates is a regression. Find what changed and revert or fix it.
- Edit effort rising means drafts are getting less useful even if they are correct.
- One category worsening points to a specific cause, such as a missing policy document.
- Noise. Models give different answers on repeated runs. If a result changes between two identical runs, run the case several times and look at the spread, not a single outcome.
Do not chase a perfect score. A set that you pass 100% of is too easy; add harder cases.
Keep it honest
- Do not train on the set. If you paste golden emails into your instructions as examples, you can no longer use them as a test. Keep a separate pool for examples.
- Rotate a few cases. Retire examples that no longer reflect your business.
- Protect privacy. Remove names, addresses, account numbers and anything confidential. Store the set with the same care as other customer data.
- Do not replace people. The golden set measures the tool. A person still reads every draft before it is sent. See piloting AI drafting with a small group for rolling it out.
A small start
If a full set feels like too much, begin with ten emails and a three-column sheet: the email, the note and a pass or fail. Add cases whenever you find a bad draft. In a month you will have a useful set, and the habit that matters.
Key takeaways
- A golden set is a fixed collection of real emails with notes on what a good reply does and avoids.
- Cover common requests, edge cases, traps and risky content, and write notes before generating drafts.
- Score correctness and safety as gates, then completeness, action, tone and edit effort.
- Re-run it after every change and add every real failure as a new case.
- Keep the set separate from your examples, private and small enough to maintain.
Start with Koltrix
Your domain, one inbox, and an API that sends.
A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.


