Auditing what an AI assistant did in your mailbox
Why every AI tool call in your mailbox should be logged, what a good audit trail records without storing mail content, a weekly review routine and red flags.

On this page(7 sections)
- Why an action log matters
- What a good audit trail records
- A weekly review routine
- 1. Check connected apps (3 minutes)
- 2. Scan activity volume (3 minutes)
- 3. Review the action types (5 minutes)
- 4. Look at denied or failed actions (2 minutes)
- 5. Note anything odd (2 minutes)
- Red flags in the log
- What to do when something looks wrong
- Choosing tools with auditability in mind
- Bottom line
Once an AI assistant can search, label, archive and draft in your mailbox, a simple question becomes surprisingly important: what did it actually do? If you can't answer that, you can't spot a misbehaving tool, a misused connection, or an assistant that followed instructions hidden in an email.
This runbook explains what an audit trail for AI mailbox actions should contain, how to review it in fifteen minutes a week, and which patterns should make you look closer.
Why an action log matters
When you do something in your own inbox, you remember it. When an assistant does something on your behalf, there's no memory to rely on, and the assistant's own chat history is not a reliable record. Conversations get deleted, summarized, or happen in a different app from the one you're looking at.
An action log gives you:
- Accountability. You can tell the difference between "I archived that" and "the assistant archived that."
- Detection. Unusual behavior, such as a burst of activity at 3 a.m. or a tool reading threads it has no reason to read, stands out in a log.
- Debugging. When a thread goes missing or gets an odd label, the log tells you how.
- Confidence for the team. People are more willing to connect assistants when they know actions are recorded and reviewable.
What a good audit trail records
The useful information is about the action, not the content of your email. A log that copies message bodies creates a second store of sensitive data that must itself be protected. A good entry looks more like this:
2026-09-18 08:42:17 UTC
user: [email protected]
app: [assistant name] (connected 2026-09-02)
action: apply_label
target: thread 7f3a… ("Invoice question")
detail: label "billing" added
result: success
At minimum, each entry should capture:
| Field | Why |
|---|---|
| Timestamp | Spot unusual times and bursts |
| User | Whose authority the action used |
| Connected app or client | Which assistant or integration acted |
| Action type | Read, search, label, archive, draft, and so on |
| Target | Which thread or mailbox, by ID or subject |
| Result | Success, denied, error |
Connection events matter as much as actions: when an app was connected, when its access was refreshed, and when it was revoked. Many of the most important questions ("Is this old integration still active?") are answered by connection history, not action history.
A weekly review routine
Fifteen minutes on a fixed day is enough for most individuals and small teams. Put it on the calendar.
1. Check connected apps (3 minutes)
List every app or assistant with access to your mailbox. For each one, ask: am I still using this? Do I recognize it? Revoke anything you don't recognize or no longer use. Unused connections are pure risk.
2. Scan activity volume (3 minutes)
Look at how many actions each app took per day. You're looking for changes, not absolute numbers. If an assistant normally does a handful of searches a day and suddenly ran hundreds, find out why.
3. Review the action types (5 minutes)
Filter to actions that change things: labels, archives, read/unread changes, drafts. Spot-check a few. Do they match what you asked for? A draft you never requested, or archiving you didn't ask for, deserves a closer look.
4. Look at denied or failed actions (2 minutes)
Denied actions are useful signals. An assistant repeatedly trying to do something it isn't permitted to do may be following instructions it picked up from an email, or a buggy integration.
5. Note anything odd (2 minutes)
Write down anything you want to follow up on. If you're on a team, share it with whoever administers the workspace.
Red flags in the log
Some patterns warrant immediate investigation:
- Activity when nobody was using the assistant, such as overnight or during a vacation.
- Reading or searching mailboxes unrelated to the user's work, especially sensitive ones like billing or HR, where the user has access but no reason to look.
- Bursts of reads immediately after an unusual inbound email. This can indicate an assistant processing a message with hidden instructions and acting on them.
- Drafts addressed to unfamiliar external recipients, particularly ones containing summaries of other threads.
- Repeated denied attempts at actions outside the app's permissions.
- An app you don't recognize, or a connection created at a time you weren't working.
None of these prove something bad happened, but each deserves a few minutes of looking.
What to do when something looks wrong
- Revoke the connection first. You can always reconnect later. Investigation is easier when the activity has stopped.
- Check the affected threads. Restore labels, unarchive, delete unwanted drafts.
- Look for the trigger. If the activity followed a specific inbound email, examine it for hidden text or instructions.
- Tell your admin if you're on a team. Other people may have the same app connected.
- Change what made it possible: narrower permissions, a different app, or a policy that sensitive mailboxes stay disconnected.
Choosing tools with auditability in mind
When evaluating an AI assistant or email integration, ask:
- Is every tool call logged, including reads?
- Can I see the log myself, or only an admin?
- Does the log avoid storing email content?
- Can I revoke a connection instantly, and does revocation take effect immediately?
- Can a workspace admin turn off AI access for everyone?
Tools that can't answer these questions clearly are harder to trust with a shared mailbox. Equally, prefer tools whose capabilities are limited in the first place. An assistant that cannot send or delete mail has a much smaller worst case, which makes the audit trail a safety net rather than your only defense.
Bottom line
If an AI assistant can act in your mailbox, you should be able to see every action it took. Log actions and connection events, not message content; review the log for fifteen minutes a week; revoke anything you don't recognize; and investigate bursts, odd hours, and denied attempts. Pair the audit trail with limited permissions so that the log is something you check, not something you rely on to undo damage.
Start with Koltrix
Your domain, one inbox, and an API that sends.
A team inbox where AI sorts and drafts (nothing is sent without your click), plus the transactional API and SMTP relay your product sends with. 7 days free, no card.


