Skip to content

Support QA: reviewing email replies without micromanaging

A lightweight support QA process: sample a few threads a week, score them on a short rubric, run calibration sessions, and give feedback that actually helps.

Koltrix Team5 min read
Clipboard, pencil and pen on a wooden surface
Photo by Kelly Sikkema on Unsplash
On this page(12 sections)
  1. The QA checklist
  2. Step 1: Choose the sample
  3. Step 2: Use a short rubric
  4. Step 3: Write feedback people can use
  5. Step 4: Deliver it in a conversation
  6. Step 5: Calibrate as a team
  7. Step 6: Rotate reviewers
  8. Step 7: Track themes, not leaderboards
  9. Reviewing AI-assisted replies
  10. What QA should not become
  11. A sample review note
  12. Key takeaways

Nobody reads every support reply their team sends, and nobody should. But if nobody reads any of them, quality drifts quietly: a wrong answer gets copied into a snippet, a curt tone becomes normal, and the first sign is a customer who cancels.

Quality review fixes that without turning a lead into a hall monitor. The trick is to review a small, regular sample against a short rubric, and to treat the results as coaching material rather than a scorecard.

The QA checklist

Here's the whole process in one list. The sections below explain each step.

  • Review a small, fixed sample per person per week
  • Pick threads randomly, plus a few chosen on purpose
  • Score with a rubric of four or five criteria, not twenty
  • Write one specific strength and one specific improvement per review
  • Share feedback within the week, in a short conversation
  • Run a calibration session every month
  • Rotate reviewers so it isn't always the lead
  • Track themes across the team, not rankings between people

Step 1: Choose the sample

You don't need volume to learn something. Three to five threads per person per week is enough for a small team to spot patterns.

Mix two kinds of selection:

  • Random picks show you what typical work looks like.
  • Deliberate picks cover the cases that matter most: escalations, refunds, angry customers, threads that were reopened, conversations with long gaps.

Avoid only reviewing threads that went badly. If every review is about a mistake, people learn to dread QA instead of using it.

Step 2: Use a short rubric

Long rubrics produce long arguments. Four criteria capture most of what matters in an email reply:

Criterion What "good" looks like Common misses
Accuracy Correct facts, policy, and steps Outdated instructions, wrong plan details
Completeness Every question in the customer's email answered Answering the first question, missing the second
Tone Clear, warm, appropriate to the situation Over-apologizing, curtness, jargon
Next step Customer knows what happens next and when No timeline, ambiguous ownership

Score each on a simple scale (meets, partly meets, misses) rather than a ten-point number. Fine-grained scores invite debates about whether something is a six or a seven, which doesn't help anyone write better emails.

Some teams add a fifth criterion for efficiency: did the reply avoid unnecessary back-and-forth by asking for everything needed up front? It's useful once the basics are solid.

Step 3: Write feedback people can use

A score alone tells someone very little. Each review should include:

  • One specific strength. "Your reply to the export question used numbered steps and linked the exact article. That's the format we want."
  • One specific improvement. "In the refund thread, the customer asked two questions and the reply covered only the refund. Try ending with a quick re-read of their email."

Specific beats general. "Be more empathetic" is hard to act on. "Acknowledge the deadline they mentioned before explaining the delay" is easy.

Step 4: Deliver it in a conversation

Written feedback is fine for small things. For anything that stings, a short call or chat works better. Ten minutes is usually enough.

A structure that works:

  1. Ask them how they felt about the thread first. People often spot their own issues.
  2. Share the strength.
  3. Share the improvement, with the specific example.
  4. Agree on one thing to try next week.

Keep it in the same week as the reply. Feedback on an email from a month ago feels like an audit.

Step 5: Calibrate as a team

Different reviewers will score the same reply differently. Calibration closes that gap.

Once a month, pick two or three threads and have everyone score them independently using the rubric. Then compare. Where scores differ, discuss why until you agree on what "meets" means. Write down any decisions ("we don't penalize short replies if all questions are answered").

Calibration sessions are also where your rubric evolves. If a criterion keeps causing confusion, rewrite it.

Step 6: Rotate reviewers

Peer review has real advantages over lead-only review:

  • Reviewers learn from reading others' replies.
  • Feedback comes from someone who does the same work.
  • The lead isn't a bottleneck.

The downside is that peers may be reluctant to criticize each other. Calibration helps, and so does framing reviews as "what I'd steal from this reply" plus "one thing I'd change".

A common split: peers review routine threads, the lead reviews escalations and sensitive cases.

Step 7: Track themes, not leaderboards

The most valuable output of QA is the list of recurring issues across the team:

  • "Three reviews this month found outdated instructions for the import flow." That's a docs or snippet problem.
  • "Several replies didn't give a timeline for bug fixes." That's a process gap: maybe support doesn't know the timelines.
  • "Tone issues cluster in billing threads." That's a training topic, or a policy that's hard to explain kindly.

Avoid turning QA scores into a ranking of individuals or tying them to compensation. Once scores affect pay, people optimize for the rubric instead of the customer, and reviewers start grading generously to avoid conflict.

Reviewing AI-assisted replies

If your team uses AI to draft replies, QA matters more, not less. The person who clicks send owns the reply, but reviewers should also note where drafts went wrong:

  • Did the reply include a detail the agent clearly didn't verify, such as a plan limit or a date?
  • Did it answer a slightly different question than the customer asked?
  • Did it use phrasing your team wouldn't, like overly formal apologies?

Track these separately from the agent's own feedback. Patterns here tell you whether the AI needs better source material or should stay out of certain categories entirely.

What QA should not become

  • A second approval step. If replies need sign-off before sending, that's a training or trust issue, not QA.
  • A blame tool. Mistakes found in QA are process signals first.
  • A secret. Everyone should know the rubric, the sample size, and how results are used.

A sample review note

Thread: Customer asking about moving data between workspaces
Accuracy: meets
Completeness: partly meets - didn't address whether labels transfer
Tone: meets
Next step: meets - clear timeline for the follow-up
Strength: Great job restating the customer's goal before answering.
Try next: End with a quick re-read of their email to catch second questions.

Key takeaways

  • Review a small, regular sample, mixing random and deliberately chosen threads.
  • Keep the rubric to four or five criteria with a simple scale.
  • Pair every review with one specific strength and one specific improvement.
  • Calibrate monthly so reviewers agree on what good looks like.
  • Use QA to find team-wide themes, not to rank individuals.

Start with Koltrix

Your domain, one inbox, and an API that sends.

A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.

SharePost on XLinkedIn