AI Inbox Agent
Classify, summarise and draft replies, with a human approving every send
An inbox agent that classifies email, summarises it and drafts replies from your past answers, with human approval. A Python CLI and an n8n + Slack version.
Solo build; two versions: a Python CLI and an n8n + Slack workflow
- 80%
- Time saved on triage · n8n + Slack version owner-reported
- 15,000+
- Emails processed · n8n + Slack version, to date owner-reported
- 92%
- Draft accuracy · n8n + Slack version owner-reported
The problem
An inbox mixes three kinds of work. Some messages need a real answer, some only need to be read, and most can be ignored. Sorting them costs a few seconds each, but those seconds add up, and the messages that need a reply are not always the ones I read first. When a reply is needed, I usually write something close to what I have written before.
I wanted a tool that does the sorting and the first draft for me, and still leaves the decision to send with a person. An auto-reply bot was never the goal. The goal was to read less noise, answer faster, and never send anything I haven't seen.
What I built
Two versions of the same idea.
The first is a Python command-line app, the one in the public repository. It works through your unread Gmail inbox with five small agent classes. For each email it asks an LLM to classify it, summarise it and, if a reply is worth sending, draft one. The draft prompt includes your three most similar past replies, retrieved from a local vector store. You approve, edit with feedback, or skip each draft in the terminal. Approved replies are sent in the original thread and saved as examples for the next draft.
The second is a workflow built in n8n with Slack. It is the version the numbers on this page come from. I'll add its n8n export or a screenshot to this page later; until then this page describes it only at that level.
The one-sentence version: the agent sorts and drafts, and a person decides what gets sent.
See it running
A recorded walkthrough of AI Inbox Agent is on the way. Until then, the write-up and the diagram below show how it works.
How it fits together
Nothing is sent until you approve it
- Gmail inboxGmail API: unread messages, once or every 5 minutesup to 50 per cyclenew email
- Classify, summaryOpenAI or Anthropic, JSON answers6 categories3 prioritiesaction itemsworth a reply
- Draft from pastLocal ChromaDB: 3 similar past repliesskips spam, promos, lowdraft
- You approveTerminal prompt on your machineapproveedit with feedbackskipapproved
- Send, log, learnReply in thread, Sheets log, new examplefollow-up reminders
Architecture
The diagram shows the Python version, because that is the code you can read.
- Fetch: the Gmail API lists unread messages in the inbox, up to 50 per cycle by default.
- Classify: an LLM returns a category (urgent, important, promotional, newsletter, spam, general), a priority (high, medium, low), a confidence score and its reasoning, as JSON.
- Summarise: a second call returns a short summary, key points, action items and sentiment.
- Decide: spam, promotional, newsletter and low-priority emails get no draft.
- Draft: the three most similar past replies come from a local ChromaDB store and go into the prompt.
- Approve: you choose approve, edit with feedback or skip.
- Send, log, learn: the reply goes out in the thread, is stored as a new example, and a row is written to a Google Sheet. Unanswered high and medium priority emails get a local follow-up reminder.
The trade-offs are deliberate. Each step is a small, single-purpose LLM call with a JSON answer, so I can change the classifier without touching the drafts. The vector store is local, so past replies never leave the machine except inside a prompt. And it polls every five minutes instead of using a push trigger, which is simpler to run on a laptop.
Stack
- Python: the CLI app and its five agent classes, coordinated by one orchestrator class.
- Gmail API with OAuth: read, send in thread, mark as read.
- OpenAI or Anthropic: one setting picks the provider; every agent asks for JSON.
- ChromaDB with sentence-transformers: a local, persistent store of past replies for retrieval.
- Google Sheets (gspread): an optional log of every processed email.
- n8n and Slack: the second version of the workflow.
Key features
- Six categories, three priorities. Every email gets a label, a confidence score and a one-line reason you can check.
- Summaries with action items. You read the summary and the actions, not the whole thread.
- Drafts from your own replies. Retrieval pulls your three closest past replies into the prompt, so drafts follow how you already answer.
- Approval by default. Approve, edit with feedback (the LLM rewrites the draft), or skip.
- It learns from what you send. Every approved reply becomes a new example.
- Follow-ups and a log. Unanswered important emails get a local reminder after three days by default, and each email is logged to a sheet.
Code or config highlight
The approval step in the Python version, from main.py (lines 102–118):
# Step 6: Human approval (if required)
if settings.reply_approval_required:
approval = self._get_human_approval(email, reply_data)
if approval == "approve":
reply_sent = self._send_reply(email, reply_data)
elif approval == "edit":
feedback = input("Enter your feedback: ").strip()
refined_reply = self.reply_drafter.refine_reply(
reply_data["body"], feedback, email
)
reply_sent = self._send_reply(email, refined_reply)
elif approval == "skip":
logger.info(f"Skipped sending reply for email {email['id']}")
else:
# Auto-send without approval
reply_sent = self._send_reply(email, reply_data)
Full file: main.py, lines 102–118.
"Edit" doesn't open an editor. Your note goes back to the LLM, which rewrites the draft, so you correct the agent in plain language. Approval is on by default; the else branch only runs if you switch REPLY_APPROVAL_REQUIRED off.
Results
These numbers are owner-reported and describe the n8n + Slack version:
- 80% time saved on triage (owner-reported).
- 15,000+ emails processed (owner-reported).
- 92% draft accuracy (owner-reported).
The Python CLI has no tracked numbers. It is a local tool you run yourself, not a hosted service, and the repository's notes record a start-up run that reached the Google sign-in step, not a measured processing run.
Lessons
Approval is the feature. The agent is useful because it drafts and stops. The moment it sends on its own, every mistake is public. That is why approval is the default in the Python version.
Plain text isn't the whole email. The Python version reads only the text/plain part, so HTML-only emails reach the agents with an empty body. Handling HTML is a fix it still needs.
Store what was actually sent. When you edit a draft, the example stored for retrieval is the original draft, not the edited reply that went out. Retrieval should learn from the final version, and that is the next change I would make.
Shared tokens need matching scopes. The Gmail client and the Sheets logger share one token file but need different scopes, so a fresh setup can fail on one of them. Separate tokens would avoid that.
Links
The Python version is in the GitHub repository, with setup steps for the Gmail OAuth client and the environment settings. You need your own Google Cloud project and an OpenAI or Anthropic key. For the n8n side of inbox triage, my n8n inbox triage tutorial walks through a build step by step. If you want the same pattern for content instead of email, see the n8n Content Machine.
Need something like this built?
I build AI automations, agents and Salesforce solutions for teams. Tell me what you want to automate and I’ll tell you honestly whether it’s a fit.
Comments
Loading comments...