All articles
·9 min

AI triage for customer emails: how to respond 3x faster without growing your team

The shared inbox: where customer requests wait the longest

Almost every B2B company has an office@ or orders@ inbox where everything lands: orders, quote requests, complaints, invoice questions, delivery confirmations and spam. Someone on the team "guards" it — usually on top of ten other responsibilities.

The numbers we see consistently at companies with 20-200 employees:

  • 80-300 emails/day on shared addresses, of which 60-70% are repetitive requests with a predictable structure
  • Median time to first response: 4-24 hours — not because the answer is hard, but because the request sits unread and unrouted
  • 30-45 minutes/day per person spent purely reading, labeling and forwarding messages that weren't meant for them
  • Between 5% and 10% of requests get lost entirely — unread, sent to the wrong person, or buried under other threads

The frustrating part: most of these emails don't need a human to be *understood*. They need a human to be *resolved* — and often not even then. That's exactly where AI triage comes in.

What an AI triage system actually does

AI triage is not "a chatbot talking to your customers". It's an invisible layer that processes every incoming message, within seconds, in four steps:

1. Classification — what kind of request it is: new order, quote request, complaint, billing question, delivery status, other

2. Extraction — the relevant data from the text and attachments: order number, requested products, customer ID, mentioned deadline

3. Routing — the message lands automatically with the right team or person, with priority set (a complaint from a top-10 customer doesn't sit in the same queue as a generic question)

4. Prepared reply — for repetitive requests, the system writes a draft based on real data from your ERP or CRM, which a human only reviews and sends

The essential difference from an autoresponder: the system doesn't reply "we have received your message". It drafts "order 4471 shipped yesterday with tracking number 123, estimated delivery Thursday" — because it read the question, identified the order and queried the courier system.

Why classic rules aren't enough

Many managers' first reaction: "Outlook rules can handle this". We've seen dozens of such rule sets. They all fail in the same place: customers don't write to a template.

A keyword rule catches "quote request" in the subject line. It doesn't catch "hello, I spoke with Mr. Ionescu on the phone, could you give us a price for 200 units of last year's model?". A language model catches both, because it understands intent, not just words.

In our tests on real-world emails — missing punctuation, half-finished sentences, multiple requests in a single message — keyword rules classify 50-65% of messages correctly. An LLM configured on the company's taxonomy reaches 90-95%. The gap between those two percentages is exactly the manual work left over.

Case study: B2B distributor with 45 employees

A technical equipment distributor we worked with at NEXVA SYSTEM received ~180 emails/day across two shared addresses. Three inside-sales people lost a combined ~3 hours/day on pure triage: reading, labeling, forwarding, looking up orders in the ERP just to be able to answer.

What we built:

  • A Microsoft 365 connector that processes every new email (including PDF attachments) in under 30 seconds
  • Classification into 11 categories defined together with the team, with a confidence threshold: below 85% certainty, the message goes to human triage, unlabeled
  • Data extraction plus automatic ERP lookups: for status questions, the draft reply already contains the tracking number and estimated date
  • Routing into team queues, with automatic escalation if a complaint sits untouched for over 4 hours
  • A dashboard with volumes, response times and trending categories — data nobody had before

Results after 3 months:

  • Median time to first useful response: from 6.5 hours to under 2 hours
  • Status requests resolved with a human-approved AI draft: 71%, with handling time under 2 minutes
  • The team's manual triage time: from ~3 hours/day to under 40 minutes/day
  • Zero lost requests in 3 months, versus an estimated 8-10/month before
  • An unplanned effect: the dashboard revealed that 22% of the volume was stock availability questions — they published stock levels in the customer portal and total volume dropped by 15%

The human stays in the loop — as a decision, not an excuse

Our firm recommendation for the first months: the AI sends nothing to customers on its own. It classifies, routes and prepares — a human approves. The reasons are practical:

  • One wrong reply to a customer costs more than the automation saves
  • Human approvals generate exactly the data you need to measure real accuracy, not demo accuracy
  • The team builds trust by watching the system work, not by being reassured

Only once a reply category consistently reaches over 95% approval without edits does auto-sending become worth discussing — and only for that category, with clear exclusions (sensitive accounts, large amounts, negative tone detected).

What it costs and when it pays back

For a company with 100-300 emails/day on shared addresses:

| Component | Cost |

|-----------|------|

| Flow analysis + taxonomy definition | €2,000-4,000 |

| Email connector + classification + extraction | €5,000-9,000 |

| ERP/CRM integration for data-backed drafts | €4,000-8,000 |

| Routing, escalation, dashboard | €3,000-5,000 |

| Monthly costs (LLM + maintenance) | €250-600/month |

Inference cost is small relative to the fear around it: at 200 emails/day, the LLM bill is under €100/month. The rest is classic engineering — connectors, integration, queues.

The ROI math is direct: 2-3 hours/day of triage saved means, at a loaded cost of €15-20/hour, €8,000-15,000/year from time alone — before putting a price on the orders that no longer get lost and the customers who get an answer the same morning instead of the next day.

The mistakes that sink triage projects

  • A taxonomy designed from a desk. Categories must come from real emails, not from the org chart. Recommended: manually classify 300-500 historical messages before writing a line of code.
  • No confidence threshold. A system forced to classify everything will fail silently. It's healthier for 10-15% of messages to go to a human than for 5% to land in the wrong queue with nobody noticing.
  • Drafts without data. A nicely worded generic reply helps no one. The value of a draft is in the data pulled from the ERP, not in the politeness of the phrasing.
  • No baseline measurement. If you don't know your current time to first response, you won't be able to prove the improvement. Measure for two weeks before you build.

How to start

1. Measure the current state: daily volume, time to first response, triage time per person

2. Manually classify a sample of 300-500 historical emails — the real taxonomy emerges from there

3. Start with classification + routing, no automated replies — immediate value, minimal risk

4. Add ERP-backed drafts for the 2-3 categories with high volume and repetitive structure

5. Expand based on numbers only — categories with proven accuracy, not the ones that merely look simple

AI triage is one of the few automations where the benefit shows up in the first week: messages reach the right person, with context, and nobody has to play dispatcher anymore. And the customer who gets an answer in two hours instead of the next day doesn't know a language model was involved — they just know you work fast.

Want to find out how many hours your team loses on manual triage and what share can be safely automated? Book a free consultation.

Want to discuss automating your processes?

Book a consultation