Skip to content

How to pilot AI drafting with a small group first

Run a two-week AI drafting pilot with two or three people: pick categories, rate edits, log failures, set go/no-go criteria, then roll out with a style guide.

Koltrix Team4 min read
Repeating AI letters on a yellow background
Photo by Steve A Johnson on Unsplash
On this page(11 sections)
  1. Why pilot at all
  2. Step 1: Pick the pilot group
  3. Step 2: Pick the categories
  4. Step 3: Prepare the inputs
  5. Step 4: Rate every draft (simply)
  6. Step 5: Collect failure examples
  7. Step 6: Check in briefly, often
  8. Step 7: Decide go or no-go
  9. Step 8: Roll out deliberately
  10. Common pilot mistakes
  11. Key takeaways

Turning on AI drafting for a whole team at once is how you end up with a week of inconsistent replies, a few embarrassing mistakes, and half the team quietly switching it off. A small pilot avoids that and tells you, with your own email, whether the drafts are worth using.

This playbook covers a two-week pilot for two or three people: how to scope it, what to record, how to decide, and how to roll out afterward.

Why pilot at all

AI drafting quality depends heavily on your mail. A tool that writes great replies to simple how-to questions may struggle with your billing edge cases, your technical customers, or your tone. Demos can't tell you that. Two weeks of real use can.

A pilot also surfaces the human questions early: who is responsible for an AI-drafted reply, how much editing is acceptable, and whether people actually trust the drafts enough to use them.

Step 1: Pick the pilot group

Two or three people is the right size. Enough to see variation, small enough to talk to everyone daily.

Good pilot members:

  • Handle real volume in the inbox every day.
  • Have different styles, for example one experienced teammate and one newer one. The experienced person spots subtle errors; the newer one shows whether drafts help someone still learning.
  • Will write things down. The pilot is only as good as its notes.

Avoid piloting only with enthusiasts. You want honest skeptics in the group too.

Step 2: Pick the categories

Don't pilot on everything. Choose two or three categories of email where you expect drafts to help and where mistakes are recoverable.

Category Good pilot candidate? Why
How-to questions Yes Frequent, factual, well documented
Bug report acknowledgments Yes Repetitive structure
Feature request replies Yes Low risk, tone matters
Billing and refunds Later Involves money and policy judgment
Cancellations Later Emotionally sensitive
Legal, security, account access No Too high-stakes to learn on

Write the chosen categories down so everyone in the pilot uses drafts on the same kinds of mail.

Step 3: Prepare the inputs

Drafts get better with context. Before day one:

  • Write a short style guide (sign-off, tone, words you avoid, things never to promise).
  • Collect a fact sheet for the pilot categories: links to help articles, current plan names, known limitations.
  • Agree on the review rule: every draft is read in full before sending, and the person who sends owns the reply.

In Koltrix, AI drafts but never sends; a person always clicks send. Whatever tool you use, keep that rule during the pilot. You are testing draft quality, not autonomous replies.

Step 4: Rate every draft (simply)

Measuring "edit distance" precisely is overkill. A three-level rating that takes two seconds is enough:

  • Minor: sent with small edits (a word, a link, a greeting).
  • Major: the structure was right but significant parts were rewritten.
  • Rewrite: discarded and written from scratch.

Keep a shared sheet with one row per drafted email:

Date | Person | Category | Rating (minor/major/rewrite) | Why (one line) | Thread link

The "why" column is the most valuable. "Invented a menu path," "too formal," "missed the actual question," "perfect." Patterns jump out after a few days.

Step 5: Collect failure examples

Separately from the ratings, keep a list of drafts that would have caused harm if sent unedited. These are the cases that matter for your go/no-go decision:

  • Wrong facts (features, prices, dates)
  • Claimed actions that hadn't happened ("I've refunded you")
  • Wrong recipient name or wrong customer context
  • Inappropriate tone for the situation
  • Anything that followed instructions written inside the customer's email

For each, note whether the reviewer caught it easily. A failure that is obvious on review is a much smaller concern than one that slips past a careful reader.

Step 6: Check in briefly, often

A fifteen-minute check-in at the end of each week works well, plus a running chat thread for "look at this one." Ask:

  • Which categories are drafts actually saving time on?
  • What are the most common reasons for major edits?
  • Has anyone stopped using drafts? Why?
  • What should be added to the style guide or fact sheet?

Update the style guide and fact sheet during the pilot. Improvement over the two weeks is a good sign.

Step 7: Decide go or no-go

Set criteria before the pilot starts, so the decision isn't just a mood. For example:

  • Go if, in the pilot categories, most drafts are rated minor, harmful failures were rare and easy to catch, and pilot members want to keep using it.
  • Extend if results improved noticeably after style-guide changes but are not there yet.
  • No-go if rewrites are common, harmful failures slip past review, or people find reviewing slower than writing.

Pick your own thresholds; the point is that they are written down in advance. Also listen to the qualitative signal. If the pilot group says "I'd be annoyed if you took this away," that's meaningful.

Step 8: Roll out deliberately

If the answer is go:

  1. Expand by category, not just by people. Turn on drafting for the pilot categories team-wide first.
  2. Share the style guide and the failure list with everyone. Showing real examples of what to watch for is the fastest training.
  3. Keep the review rule. Every draft read in full, sender owns it.
  4. Pick the next categories to trial, perhaps billing, using the same rating method with a smaller group.
  5. Revisit in a month. Are people still using drafts? Have new failure types appeared?

Common pilot mistakes

  • Too broad. Piloting on all mail makes results impossible to interpret.
  • No baseline expectations. Without criteria, the decision becomes whoever argues loudest.
  • Only rating the good drafts. People forget to log the rewrites. Make rating part of sending.
  • Ignoring the reviewer's time. If checking a draft takes as long as writing a reply, the tool hasn't helped yet.

Key takeaways

  • Pilot with two or three people for about two weeks on a few low-risk categories.
  • Rate drafts as minor, major, or rewrite, with a one-line reason.
  • Track harmful failures separately and note whether review caught them.
  • Write go/no-go criteria before you start.
  • Roll out by category, keep human review before send, and share the failure list as training.

Start with Koltrix

Your domain, one inbox, and an API that sends.

A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

SharePost on XLinkedIn
More in Guides →