AI Engineering13 min readAugust 21, 2026

Document AI for Invoices, Purchase Orders & Statements: Enterprise Architecture Lessons from 80K+ Docs/Month

How I architected an enterprise Intelligent Document Processing (IDP) platform extracting structured data from 80,000+ complex invoices and purchase orders per month using Claude Sonnet 5, Python, and Human-in-the-Loop workflows.

Document AIInvoice OCRPurchase Order OCRIntelligent Document ProcessingClaude Sonnet OCRStructured Data Extraction

Enterprise Intelligent Document Processing Pipeline Architecture

For decades, enterprise document extraction relied on traditional Optical Character Recognition (OCR) tools like Tesseract, Abbyy, AWS Textract, and template-based parsers like Nanonets.

In production, these legacy systems consistently broke down when exposed to real-world complexity:

  • Multi-page invoices with tables split across page boundaries.
  • Handwritten corrections, vendor stamps, and skewed mobile photos.
  • Unpredictable tax representations and rental discount calculations.
  • Layout mutations whenever a vendor updated their invoice design.

In my work at Venture7 Technologies, I co-architected and deployed an Intelligent Document Processing (IDP) platform for a US-based enterprise client, replacing their brittle legacy Nanonets pipeline with a Claude Sonnet 5 vision reasoning architecture targeting ~80,000 documents per month (invoices, purchase orders, statements, and equipment rental slips).

Here is the full architectural breakdown, engineering challenges, and lessons learned from deploying Document AI at enterprise scale.


1. High-Level System Architecture

┌─────────────────────────┐
│ Multi-Source Ingestion  │ (Email attachments, S3 buckets, Angular UI uploads)
└────────────┬────────────┘
             │
             ▼
┌─────────────────────────┐
│   Python Pre-Processor  │ (PDF to normalized PNG, DPI enhancement, orientation fix)
└────────────┬────────────┘
             │
             ▼
┌─────────────────────────┐
│ Dynamic Seller Profiler │ (Identifies vendor & retrieves vendor-specific Markdown rules)
└────────────┬────────────┘
             │
             ▼
┌─────────────────────────┐
│ Multimodal AI Extractor │ (Claude Sonnet 5 with spatial reasoning & structured JSON schema)
└────────────┬────────────┘
             │
             ▼
┌─────────────────────────┐
│ Post-Processing Engine  │ (Deterministic math verification, line-item reconciliation, regex)
└────────────┬────────────┘
             │
             ▼
┌─────────────────────────┐
│   Confidence Scoring    │
└────────────┬────────────┘
             │
    ┌────────┴────────┐
    │                 │
    ▼ (Score >= 0.88) ▼ (Score < 0.88)
┌────────────────┐ ┌─────────────────────────┐
│ Auto-Export to │ │ Human-in-the-Loop (HITL)│
│ ERP/Salesforce │ │ Next.js Review Queue    │
└────────────────┘ └─────────────────────────┘

2. Core Architectural Components

A. Pre-Processing & Image Normalization

Before sending a document to an LLM, optical preprocessing drastically improves extraction fidelity and minimizes token consumption:

  • Multi-page PDFs are split into high-resolution PNG pages (300 DPI).
  • Deskewing and contrast normalization algorithms clean background artifacts and remove scanner shadows.
  • Fast local OCR runs as a lightweight fallback for plain text headers.

B. Dynamic Seller-Specific Markdown Reasoning Rules

In enterprise accounting, every vendor formats data differently. ACME Corp includes freight inside line items, while Global Equipment lists rental tax as a separate footer line.

Hardcoding these edge cases into a monolithic prompt causes prompt bloat and degrades overall model performance. Instead, we built a dynamic seller indexing architecture:

  1. Fast header analysis identifies the vendor name or tax ID.
  2. The system loads a concise, targeted Markdown rule profile specifically for that vendor.
  3. The rules are dynamically injected into Claude Sonnet's context window alongside the image.
<!-- Sample vendor-specific rule: sellers/cat_rental.md -->
### Vendor Profile: Caterpillar Rental Services
- **Invoice Number Location:** Top-right corner prefixed with "RNT-"
- **Line Items:** Equipment rental days must be multiplied by Daily Rate. Ignore environmental recovery fee in subtotal; parse as separate surcharge.
- **Taxes:** State and Local taxes are reported separately. Sum both into `tax_amount`.

C. Multimodal Vision Extraction with Claude Sonnet 5

We chose Claude Sonnet over other LLMs due to its unmatched spatial layout reasoning and strict JSON compliance on multi-column financial documents.

# Production Document AI Extraction Microservice
import anthropic
import json
from pydantic import BaseModel, Field

class LineItem(BaseModel):
    description: str
    quantity: float
    unit_price: float
    total_price: float

class ExtractedInvoice(BaseModel):
    invoice_number: str
    invoice_date: str
    vendor_name: str
    subtotal: float
    tax_amount: float
    total_amount: float
    line_items: list[LineItem]
    confidence_rating: float = Field(description="Self-evaluated extraction confidence 0.0 to 1.0")

def run_document_ai_extraction(image_b64: str, vendor_rules: str) -> dict:
    client = anthropic.Anthropic()
    
    prompt = f"""You are an enterprise document intelligence parser.
Examine this invoice image and extract all structured data according to the schema.

VENDOR-SPECIFIC RULES:
{vendor_rules}

Instructions:
1. Verify line item sums: quantity * unit_price == total_price.
2. Ensure subtotal + tax_amount == total_amount.
3. Output strictly valid JSON matching the schema."""

    response = client.messages.create(
        model="claude-3-5-sonnet-20241022",
        max_tokens=3500,
        temperature=0.0,
        messages=[{
            "role": "user",
            "content": [
                {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": image_b64}},
                {"type": "text", "text": prompt}
            ]
        }]
    )
    
    return json.loads(response.content[0].text)

3. The 3 Non-Negotiable Safeguards for Production IDP

When dealing with financial transactions, a 1% error rate can result in thousands of dollars in accounting discrepancies. We enforced three critical validation layers:

1. Deterministic Math & Cross-Verification

Never trust the LLM's internal arithmetic. After JSON extraction, a Python validator re-calculates: $$\sum (\text{Line Item Totals}) \stackrel{?}{=} \text{Subtotal}$$ $$\text{Subtotal} + \text{Tax} + \text{Shipping} - \text{Discounts} \stackrel{?}{=} \text{Total Amount}$$

If the calculated sum deviates from the parsed total by more than $$0.02$, the document is immediately flagged for review.

2. Confidence-Based HITL (Human-in-the-Loop) Queue

  • Documents with $\ge 90%$ confidence and verified math pass automatically to downstream ERPs (Microsoft Dynamics 365 Business Central / Salesforce).
  • Documents with missing fields, unrecognized vendors, or math discrepancies are routed to an internal review portal built with Next.js 15. Operators can review bounding-box overlays and fix fields in under 10 seconds.

3. Downstream ERP & CRM Synchronization

Structured records are pushed via transactional REST APIs into the client's enterprise systems with full audit logging, raw document snapshots, and timestamped extraction metadata.


4. Production Results & ROI

After rolling out this Document AI platform:

  • Throughput: ~80,000 documents processed monthly.
  • Auto-Processing Rate: 92% of standard documents processed end-to-end without human intervention.
  • Cost Reduction: Cut third-party per-page SaaS licensing fees by over 65% compared to legacy OCR platforms.
  • Cycle Time: Reduced invoice turnaround from 48 hours to under 20 seconds.

Build a Custom Document AI Platform for Your Organization

Whether you are processing invoices, bills of lading, medical records, bank statements, or legal contracts, off-the-shelf software often fails to adapt to your unique document formats.

I design, build, and deploy custom Intelligent Document Processing (IDP) and Document AI pipelines tailored to your exact business rules.

Learn more about my AI engineering services, inspect my IDP case study, or reach out directly on WhatsApp to discuss your document automation use case.

Written by

Nikhil Nishad

AI Engineer & Freelance Full Stack Developer at Venture7 Technologies. Building enterprise document intelligence, autonomous AI workflows, and high-performance Next.js 15 web apps.

Related Articles