Team inbox metrics worth tracking, and the ones to ignore
Which shared inbox numbers tell you something real, which ones mislead, and how each can be gamed. A decision table for small teams measuring email work.

On this page(7 sections)
- Start with the question, not the metric
- The decision table
- The metrics worth tracking
- Time to first reply, median and 90th percentile
- Threads waiting longer than N hours
- Reopen rate
- Volume by category
- The metrics to use with caution
- The metrics to ignore
- How to review metrics without harming the team
- Key takeaways
Put a number on a dashboard and people will move it, whether or not moving it helps anyone. That's the core problem with measuring a team inbox: the easy numbers are often the wrong ones, and the wrong ones change behavior in ways you didn't intend.
This post sorts common inbox metrics into three groups: worth tracking, useful with caution, and best ignored. For each, you'll see what it actually tells you and how it tends to get gamed.
Start with the question, not the metric
Before choosing any number, decide what you want to know. For most small teams, it's one of these:
- Are customers waiting too long for a first answer?
- Are things falling through the cracks?
- Are we actually solving problems, or just replying?
- Is the workload growing, and where?
Every metric below maps to one of these questions. If a number doesn't answer a question you care about, don't track it, even if your tool shows it for free.
The decision table
| Metric | What it tells you | How it gets gamed | Verdict |
|---|---|---|---|
| Time to first reply (median) | Typical wait before a human responds | Empty "we got it" replies | Track |
| Time to first reply (90th percentile) | How bad the worst waits are | Same, plus cherry-picking easy threads | Track |
| Threads waiting longer than N hours | What's falling through right now | Closing threads prematurely | Track |
| Reopen rate | Whether replies actually solved the problem | Discouraging customers from replying | Track |
| Volume by category | Where work comes from and how it shifts | Mislabeling to make categories look smaller | Track |
| Time to resolution | Total time to solve, not just respond | Closing early, splitting threads | Use with caution |
| Replies per thread | Whether problems take many rounds | Writing longer single replies that don't help | Use with caution |
| CSAT on email threads | How customers felt | Only surveying happy outcomes | Use with caution |
| Emails sent per person | Activity, not outcomes | More, shorter, less useful replies | Ignore |
| Average handle time | Time spent per thread | Rushing hard threads | Ignore |
| Inbox count at end of day | Whether the inbox looks empty | Archiving without replying | Ignore |
The metrics worth tracking
Time to first reply, median and 90th percentile
The median tells you the typical experience. The 90th percentile tells you about the customers who waited longest, which is where complaints come from. Averages hide both: one email that waited four days can drag an average way up, while a pile of fast replies can hide several slow ones.
Measure in business hours if you only staff business hours, and say so. A Friday evening email answered Monday morning shouldn't look like a failure if your published hours say weekdays.
Guard against gaming: don't count automated acknowledgments as a first reply. Only a human response counts.
Threads waiting longer than N hours
This is the single most actionable number for a small team. Pick a threshold that matches your promise to customers (for example, eight business hours) and look at how many threads are over it right now.
It's a live number, not a historical one, which is the point. It tells you what to do in the next hour.
Guard against gaming: watch for threads closed without a reply. A quick spot-check of recently closed threads once a week catches this.
Reopen rate
When a customer replies to a thread you thought was done, something wasn't solved, wasn't clear, or wasn't complete. A rising reopen rate is an early sign of rushed answers.
Guard against gaming: never discourage customers from replying. Phrases like "no need to reply" in closing messages can quietly push this number down while hiding real problems.
Volume by category
How many threads arrive per week in each category: billing, bugs, how-to, sales, and so on. This is the metric that tells you whether to write better docs, fix a confusing feature, or hire.
Guard against gaming: categories only mean something if they're applied consistently. Keep the list short and the definitions written down.
The metrics to use with caution
Time to resolution sounds ideal but is messy. Some threads are legitimately long (a bug waiting on engineering), and resolution depends on when someone marks a thread done. Look at it by category, not as a single team number.
Replies per thread can reveal unclear first answers, but a complex technical issue legitimately takes several rounds. Use it to find patterns, not to judge individuals.
CSAT is useful when response rates are reasonable and you read the comments, not just the score. On its own, it tends to reflect the outcome (did they get what they wanted?) more than the quality of the reply.
The metrics to ignore
Emails sent per person rewards activity. The fastest way to increase it is to write shorter, less complete replies that require more back-and-forth. It also punishes whoever takes the hard threads.
Average handle time pushes people to rush the threads that most deserve care. It's borrowed from call centers, where calls are more uniform than email.
End-of-day inbox count measures how the inbox looks rather than whether customers were helped. Archiving everything gets it to zero.
How to review metrics without harming the team
- Review team numbers, not individual leaderboards. Individual metrics on a public screen create competition over the wrong things.
- Pair every number with a few real threads. A rise in first-reply time means more when you look at the actual emails that waited.
- Change one thing at a time. If you change the triage process and staffing in the same week, you won't know which helped.
- Revisit the list quarterly. Drop metrics nobody acted on.
A simple weekly view for a small team can fit on one screen:
Week of [date]
First reply, median (business hrs): [x]
First reply, 90th percentile: [x]
Waiting > 8 business hrs right now: [x]
Reopen rate: [x]
Volume by category: billing [x] · bugs [x] · how-to [x] · sales [x]
Notes: [what changed, what we'll try]
Key takeaways
- Choose metrics by the question they answer, not because your tool displays them.
- Track first-reply time (median and 90th percentile), threads waiting too long, reopen rate and volume by category.
- Treat resolution time, replies per thread and CSAT as context, not targets.
- Ignore activity metrics like emails sent and handle time; they reward the wrong behavior.
- Review numbers as a team, alongside real threads.
Start with Koltrix
Your domain, one inbox, and an API that sends.
A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.


