Last year we deployed a document processing system for a large logistics company that handles about 8,000 invoices per month from 340 different suppliers. After 12 months in production and roughly 100,000 processed invoices, here's what we actually learned — not from benchmarks, but from watching the thing run in the real world.

The first week was humbling

Our test accuracy was 94%. On real invoices from real suppliers, week one accuracy was 71%.

The gap wasn't the model's fault. It was our assumptions. We had trained on a well-curated dataset. Real invoices came with:

  • Scanned documents where the scanner bed was dirty (consistent gray smudges in the same spot on every page)
  • Handwritten corrections in pen over printed text
  • Three suppliers who sent invoices as Excel spreadsheets embedded in Word documents, embedded in emails, attached as .msg files
  • One supplier who typed invoice amounts in the email body and attached a blank PDF "for your records"

Nobody tells you about the .msg-inside-Word-inside-Excel pipeline in the AI marketing materials.

The 80% that's easy and the 20% that takes all your time

After two weeks of tuning, we got to 89% fully automated processing. Getting from 89% to 93% took another month. Getting from 93% to 96% took three months.

This is the dirty secret of document automation: the last 10% of accuracy improvement takes 80% of the effort. And for most businesses, 90% automation with clean human handoff is far more valuable than 99% automation that occasionally makes confident mistakes.

We restructured the whole system around this insight. Instead of trying to push accuracy to 99%, we built three tiers:

Tier 1 — High confidence (78% of invoices): Fully automated. System extracts data, validates against PO, and books it. No human touches it.

Tier 2 — Medium confidence (15% of invoices): System extracts data but flags specific fields it's uncertain about. A human reviews only the flagged fields, not the entire document. Average review time: 45 seconds.

Tier 3 — Low confidence (7% of invoices): Routed to a human for full manual processing. But the system still pre-fills what it can, saving about 60% of the manual work.

The result: 78% fully automated, 15% semi-automated, 7% manual. The team that used to spend 3 days on invoices now spends about 4 hours, mostly on Tier 2 reviews.

What actually broke in production

Month 2: A major supplier changed their invoice template. No warning. Our extraction rules for their specific format stopped working overnight. We had 400 invoices pile up before anyone noticed. Fix: We added template change detection that alerts the team when a supplier's invoice format deviates significantly from the known pattern.

Month 4: The scanning hardware was replaced with a newer model. The new scanner produced slightly different color profiles and DPI. Our pre-processing pipeline was tuned for the old scanner's characteristics. Accuracy dropped 8% for two weeks until we re-calibrated.

Month 7: We discovered that 12% of "correct" extractions had a subtle bug: amounts in certain currencies were being parsed with the wrong decimal separator. Hungarian invoices use commas as decimal separators. Some suppliers used periods. Our system was inconsistent about which convention it expected. This had been silently wrong since launch. The total financial impact was about €23,000 in mis-booked amounts.

That last one hurt. We now have a reconciliation check that runs nightly and compares extracted amounts against bank statement data.

The supplier problem

Nobody talks about this enough: your accuracy varies wildly by supplier. We tracked per-supplier accuracy for all 340 suppliers. The range was:

  • Best: 99.2% (large suppliers with consistent, machine-generated invoices)
  • Worst: 34% (a small family business that handwrites invoices on pre-printed forms)
  • Median: 91%

The bottom 20 suppliers (by accuracy) generated 60% of all exceptions. We considered three approaches:

1. Fine-tune a model specifically for the worst suppliers (expensive, needs continuous updating)

2. Route those suppliers directly to manual processing (simple, but defeats the purpose)

3. Ask the suppliers to switch to a standard electronic format (surprisingly effective — 8 of the 20 agreed within a month when we explained the benefit to both parties)

We ended up doing all three, selectively. The key insight: don't try to solve every document format with AI. Sometimes the right solution is fixing the upstream problem.

Metrics that actually matter

After a year, here's what we track (and what we stopped tracking):

What we track:

  • Straight-through processing rate: 78% (the percentage of invoices that need zero human intervention)
  • Average human review time per invoice: 45 seconds for Tier 2, 4 minutes for Tier 3
  • Time to detect template changes: Under 24 hours (was initially days)
  • Financial accuracy: Reconciliation error rate under 0.1% (after the Month 7 fix)
  • Per-supplier accuracy trends: Monthly, with automatic alerts for drops

What we stopped tracking:

  • Overall model accuracy percentage: It's a vanity metric that hides the reality of per-supplier variation
  • Processing speed: The model is fast enough. Speed was never the bottleneck. Human review queue management was.

Would we do it again?

Yes, without hesitation. The system pays for itself in about 2.5 months of operation. But we would do several things differently:

1. Start with a smaller supplier set. We tried to handle all 340 suppliers from day one. We should have started with the top 50 (which represented 80% of invoice volume) and added others gradually.

2. Build the reconciliation check from day one. The Month 7 decimal separator bug would have been caught in week one.

3. Set client expectations around the tiered approach upfront. The client initially expected 100% automation. Managing that expectation early would have avoided friction.

4. Budget more for the "long tail" of edge cases. The 7% in Tier 3 aren't getting automated anytime soon. That's okay. The ROI is already massive.

If you're considering document automation, this is what it actually looks like. Not the clean demo with 10 sample invoices — the reality of 8,000 invoices from 340 suppliers, every month, for a year.

The step-by-step decision guide is in AI Document Processing in Practice.