AI & AutomationStructured JSON Output · 90%+ Field Accuracy · Client POC DeliveredPrivate Codebase

AI OCR Document Intelligence Platform

Enterprise OCR + LLM parsing pipeline for structured invoice & receipt data extraction

AI OCR Multi-Engine Pipeline 3D mascot - cyber detective robot with antenna and high-tech laser magnifying glass

Technologies Used

PythonNode.jsOCR EngineVision LLMsLangChainMongoDBFastAPIJSON Schema Validation

Proprietary / Private Codebase

To protect client confidentiality, proprietary AI models, and commercial IP, this repository is private. If you are an engineering manager, recruiter, or collaborator wishing to explore code samples, architecture specifics, or arrange a private code walk-through, feel free to connect directly:

Executive Overview

An enterprise document intelligence platform that extracts key tabular and line-item data from invoices, purchase orders, and receipts — combining computer vision OCR with LLM post-processing to output validated, schema-enforced JSON. Delivered as a client engagement to automate manual data entry workflows for a logistics business processing 500+ invoices monthly.

The Problem Statement

A logistics client was spending 40+ staff-hours monthly manually transcribing invoice line items (quantities, unit prices, totals, vendor codes) from scanned PDFs into their ERP system. The error rate from manual entry was causing reconciliation failures downstream, costing the team additional correction time weekly.

The Engineering Solution

Built a two-stage document processing pipeline: Stage 1 — OpenCV preprocessing (denoising, deskewing, contrast normalization) feeds a Tesseract OCR engine to extract raw text from invoice scans. Stage 2 — The raw OCR text is passed to an LLM (GPT-4o) with a strict function-calling schema to semantically parse and validate line items, totals, vendor info, and date fields into a structured JSON output. Node.js backend queues process invoices asynchronously with MongoDB storing extracted results.

System Architecture & Implementation

Python OCR microservice (FastAPI) handles image preprocessing and Tesseract extraction. Node.js Express backend orchestrates the queue and calls the Python service, then passes OCR text to the OpenAI function-calling API for structured extraction. MongoDB stores raw OCR output alongside validated JSON for audit trails. JSON Schema validation layer rejects outputs with missing required fields before delivering to the client API endpoint.

Technical Challenges & Solutions

Handling low-contrast fax-quality document scans and varied multi-vendor invoice layouts (table alignments differ per supplier). Some invoices mixed printed and handwritten annotations, causing raw OCR to produce garbled text mid-field. Getting the LLM to produce consistent, schema-validated JSON without hallucinating prices or quantities required 15+ prompt iterations with chain-of-thought reasoning and explicit examples.

Key Lessons Learned

Using LLMs as a secondary semantic parsing layer after raw OCR text extraction dramatically improves structured field accuracy — from ~65% with regex-only parsing to ~90%+ with LLM post-processing. The function-calling API approach was more reliable than free-form JSON requests because it enforces the output schema at the API level, eliminating parsing errors.

Roadmap & Future Improvements

Supporting multi-page PDF batch processing with automated validation webhooks that notify the client system only on high-confidence extractions, routing low-confidence results to a human review queue.

Case Study Narrative & Verified Outcomes

Deep-Dive Engineering Breakdown

Modern enterprise document workflows routinely choke on multi-vendor invoice formats. While traditional OCR solutions extract raw character text, they fail to reconstruct tabular relationships—such as associating an item code with its corresponding unit price and quantity when descriptions span multiple lines.

To solve this for our logistics client, we engineered a decoupled hybrid vision + semantic processing pipeline:

[ Scanned Invoice (PDF/TIFF) ]
              │
              ▼
    ┌──────────────────┐
    │ OpenCV & Deskew  │ ──► Denoising & Otsu Binarization (300 DPI)
    └─────────┬────────┘
              │
              ▼
    ┌──────────────────┐
    │ Tesseract Engine │ ──► Spatial Bounding Box Token Generation
    └─────────┬────────┘
              │
              ▼
    ┌──────────────────┐
    │ FastAPI Service  │ ──► Coordinate-Based Table Reconstruction
    └─────────┬────────┘
              │
              ▼
    ┌──────────────────┐
    │ GPT-4o Function  │ ──► Strict Pydantic JSON Schema Validation
    └─────────┬────────┘
              │
              ▼
[ Validated Accounting JSON ──► Automated ERP Database Sync ]

1. Image Preprocessing & Contrast Normalization

Low-quality mobile photographs and faxed receipts routinely suffer from shadows, skew, and uneven contrast. Using OpenCV in Python, we applied:

  • Hough Line Transform to detect baseline angles and deskew scans up to 45 degrees.
  • Bilateral Filtering to smooth background noise while preserving crisp character edges.
  • Adaptive Thresholding to normalize white balances across faded paper sheets.

2. Strict Schema Validation & Mathematical Cross-Footing

A common failure mode in LLM document parsing is hallucinations on numerical fields. To eliminate this:

  • We enforced Strict Structured Outputs using Pydantic models.
  • Implemented a deterministic Python mathematical validation check: quantity * unit_price == line_total and sum(line_totals) + tax == grand_total.
  • If the validation fails by more than a 2-cent rounding margin, the document is automatically flagged and routed to a Human-in-the-Loop review queue.

3. Real-World Business Outcomes

  • Time Savings: Reduced manual data entry from 40 staff-hours to under 2 hours of human verification monthly.
  • Accuracy: Reached 90%+ straight-through field accuracy on standard printed invoices.
  • Turnaround: Cut document-to-ERP latency from 3 business days down to under 25 seconds per document.

Related Reading: Learn more about our AI automation engineering services or read our deep-dive guide on production invoice OCR architectures.

Related Engineering Case Studies

Companion Engineering Guide
All 30 Articles

Document AI for Invoices, Purchase Orders & Statements (80K+ Docs/Month)

Deep-dive into the Claude Sonnet 5 + Python enterprise extraction pipeline architecture.

Interested in similar architecture?

Work With Nikhil Nishad

I partner with startups, product teams, and founders globally to design, build, and deploy production-grade AI systems, Next.js web applications, and document automation.