A deep dive into building document intelligence systems that understand context, relationships, and the difference between extracting text and extracting meaning.
I'm knee-deep in OCR right now, and let me tell you, there's a world of difference between extracting text and extracting intelligence. After three weeks of building what started as a simple PDF-to-search pipeline, I've ended up with something far more sophisticated: a contract intelligence platform that can answer questions like "What are the current payment terms for Client X?" while respecting the full hierarchy of master agreements, addendums, and amendments.
The journey from basic OCR to true document intelligence revealed fundamental insights about LLM-based extraction that I wish someone had told me at the start. Most importantly: Schema = Instructions. The LLM only extracts what you explicitly ask for, nothing more. This single realization drove everything that followed.
Picture this: thousands of pages of legal contracts that are scanned, faint, multi-generation fax copies with inconsistent layouts. Account managers need to answer high-stakes questions about feature deprecations, module end-of-life, and customer commitments. A missed clause can mean financial penalties or compliance violations.
The current process? Manual PDF searching, institutional memory, and crossed fingers.

Simple OCR wasn't going to cut it. Neither was throwing everything into a vector database and hoping semantic search would figure it out.
THE OPTIONS: I tested five different approaches, each with distinct strengths for different parts of the document processing pipeline.

What makes it different: Single API for both parse and extract with native JSON schema support.
Why I chose it: LandingAI became my primary platform because it's purpose-built for document extraction workflows. You can parse a PDF to markdown, then run structured extraction with custom schemas in one unified pipeline. The markdown output preserves document structure naturally, and you get grounding information (bounding boxes) for citations.
Performance notes: They handle degraded scans remarkably well. I threw everything from faint faxes to rotated pages at it, and it consistently outperformed general-purpose solutions.
Trade-offs: Per-page credit pricing scales with volume, but the unified workflow saves significant development time compared to stitching together multiple services.
What makes it different: Native image processing with strong layout understanding.
Why I considered it: Gemini excels at understanding document structure. It can process images as first-class inputs and has excellent printed text accuracy. For general OCR tasks, it's impressive.
Performance notes: Good for general OCR and document Q&A where you don't need structured outputs.
Trade-offs: No structured extraction API. You have to prompt for structure, which is non-deterministic and expensive. I implemented repetition detection because temperature matters more for extraction accuracy than I expected.
What makes it different: Best-in-class for forms and tables.
Why I use it: Textract is the go-to for structured layouts. If you have pricing sheets, order forms, or expense documents, this is your tool. The table extraction is native and cell-level accurate.
Performance notes: Excellent for anything with clear tabular structure—pricing sheets, schedules, order forms.
Trade-offs: Weaker on dense paragraphs and legal prose. Higher cost at scale compared to other options.
What makes it different: PDF-to-markdown conversion optimized for reading order.
Why it surprised me: Marker achieved a 96.69% heuristic score on legal documents and excels at reading order detection—crucial for multi-column layouts. It works entirely locally (no API costs) and is fast (~0.18s/page on GPU).
Performance notes: Best for layout reconstruction when you need to preserve document structure for downstream processing.
Trade-offs: Not a top-tier OCR engine for degraded scans. It relies on PDF text layers, so scanned documents need preprocessing with actual OCR first.
What makes it different: It's the proven, lightweight option that's been around forever.
Why I use it: Tesseract provides fast baseline extraction and good character count benchmarking. It's what I use for sanity checks and coverage comparison during development.
Performance notes: Fast and lightweight for development testing and baseline metrics.
Trade-offs: Weak on complex layouts and poor confidence scoring. Not suitable for production contract extraction but valuable for testing.
THE TAKEAWAY: Multi-model OCR isn't about finding one perfect tool—it's about knowing which tool solves which problem and building intelligent routing between them.
The real breakthrough came when I realized that static schemas fundamentally don't work for diverse document types. Here's the journey that led to that insight.

My initial approach was textbook: create one comprehensive schema that extracts everything from any contract type. I built a RAG-optimized schema with clauses, risk levels, keywords, and source pages. Testing it on a service provider agreement worked beautifully:
I was feeling pretty good about this until I tested it on a different document—a Progress Software Addendum containing pricing tables, 14 partner airlines with passenger numbers, and license schedules.
Static schema results: 13 clauses, 0 tables, 0 partner airlines, 0 license schedules.
The LLM did exactly what I asked for: clauses. It completely ignored the structured data because my schema didn't mention tables.
That's when it hit me: Schema is not just configuration—it's instruction. The LLM extracts exactly what you ask for, nothing more. If you don't ask for pricing tables, you won't get pricing tables, even if they're the most important part of the document.
This led to the obvious-in-hindsight question: "How do we know what to extract from a document we haven't seen yet?"
The answer was a three-phase pipeline:
Convert PDF to markdown using LandingAI Parse
LLM analyzes structure, identifies tables/lists/entities
Generate schema from discovery, run extraction
Testing the dynamic approach on the same Progress Software Addendum:
The extraction phase actually found more than discovery because the generated schema prompted the LLM to find additional instances of discovered patterns.
Cost trade-off: 9 credits vs 6 for static (50% more), but dramatically better results. For contract intelligence, accuracy matters more than cost savings.
Traditional RAG hits a fundamental wall with contract hierarchies. Real contracts don't exist in isolation—you have master agreements, addendums, amendments, and supersessions. When someone asks "What are the payment terms for Client X?", they need the current terms, not the original ones that were changed by Addendum B six months ago.
Traditional RAG treats each document independently. It can't answer "What's current?" because it doesn't understand relationships.

Consider this structure:
(Net 30)
(Net 45)
(Net 60)
(2-year extension)
When querying payment terms, you need:
I implemented a knowledge graph that explicitly tracks relationships. The graph stores clients, contracts, clauses, and terms with explicit relationships. Clauses are categorized (payment_terms, liability, termination) and embedded for semantic search. The system can traverse the graph to find the most recent effective term for any category.
Key queries this enables:
The graph auto-populates from extraction results using pattern detection:
This caught supersessions that manual review would miss, increasing detected relationships from 1 to 18 in our test corpus.
The GraphRAG service evolved from a 4,215-line monolith into a modular architecture with 9 focused modules:
graph_rag/
├── __init__.py # Public API composition
├── models.py # Data models
├── utils.py # Helpers
├── core.py # CRUD operations
├── persistence.py # Save/load
├── query_engine.py # Queries & search
├── import_service.py # Import & parsing
├── linkage_detector.py # Auto-linking algorithms
├── cross_contract_query.py # Portfolio-wide queries
└── db_repository.py # PostgreSQL storageThis refactoring reduced the largest module by 52% and improved developer onboarding from 3 days to 1 day.
What started as a local development setup needed to become a shared deployment. The architecture evolution tells an important story about scaling document intelligence systems.

Initially, I had data scattered across six different storage mechanisms:
This works fine for single-instance development but breaks completely in multi-instance deployments.
The solution was consolidating everything possible into PostgreSQL with pgvector. I completed all four phases of migration:
Moved 50 clients, 111 contracts, and 950 clauses from JSON to PostgreSQL tables with proper relationships and foreign keys.
Migrated 82,180 chunks with embeddings to pgvector, enabling native similarity search. I implemented dual embedding columns—1536 dimensions for OpenAI's text-embedding-3-large and 768 dimensions for Google's embeddings—to support multiple providers.
Added a cache layer for RAG query results with 1-hour TTL. Sub-millisecond cache lookups with cache invalidation on document re-indexing. The system includes management endpoints for cache statistics and selective invalidation.
Combined pgvector semantic search with PostgreSQL full-text search using Reciprocal Rank Fusion (RRF) scoring:
Default weights: 70% semantic, 30% keyword. This catches both semantic matches and exact terminology that semantic search might miss. Each result returns both semantic_score and keyword_score for transparency.
Testing revealed surprising cases where traditional RAG struggled compared to direct LLM queries with full context.

"What are the termination clauses across all contracts?"
"I cannot find this information in the retrieved sections. The context provided only includes repeated fragments..."
Correctly identifies and explains all termination provisions across documents.
Root Cause:
Poor chunking creates silent failures. When a clause gets split into fragments like "tomers.", "ther customers.", "mers.", these fragments don't carry semantic meaning for search, but they show up in results and waste context window space.
Solutions implemented:
Never split mid-clause
In continuation chunks
Enforcement
For context preservation
The hybrid strategy: Start with RAG for efficiency, fall back to direct queries for complex cases.
The final architecture reflects lessons learned about what works in production document intelligence systems:

~40 service files organized by concern, 17 API router groups, 14 database models. The modular structure emerged from refactoring a 4,215-line monolithic service into focused components.
Key services:
Handles all chunking strategies (smart_rag, clause_boundary, simple)
PostgreSQL-backed vector operations with caching
Modular composition using multiple inheritance
Provider abstraction with fallback chains
Five main interfaces reflecting the workflow:
PostgreSQL handles everything:
The unified storage approach eliminates sync issues and enables complex queries across data types.
Building a working prototype is one thing. Getting it into the hands of users who actually need it—account managers answering client questions, sales teams preparing for renewals—is another challenge entirely.

The platform runs on AWS with a production-grade setup:
The ALB routes /api/* to the backend service and everything else to the frontend—simple but effective for this architecture.
For an internal tool handling sensitive contract data, security was non-negotiable:
Cognito-based authentication integrated with the load balancer
API key authentication for system-to-system integration
Database and cache in private subnets with no public access
TLS everywhere, secrets managed outside the codebase
A few gotchas that cost me time:
Platform mismatch: Docker images built on Mac (ARM64) won't run on Fargate (AMD64). Always specify --platform linux/amd64 when building for cloud deployment.
pgvector installation: RDS PostgreSQL doesn't auto-enable extensions. You have to explicitly run CREATE EXTENSION vector from within the VPC—which means spinning up a temporary container just to run that command if your database isn't publicly accessible.
Nginx routing: The frontend container tried to proxy API requests to a hostname that doesn't resolve in ECS networking. Solution: let the load balancer handle all routing, don't duplicate it in nginx.

Here's the thing about enterprise tools: nobody wants another application to check. Account managers live in email and chat. Sales lives in their CRM. Asking people to context-switch to a new web UI is asking for low adoption.
So I integrated the platform with BrainTrust, our internal agentic AI system. BrainTrust puts an AI agent directly in Google Chat—where our teams already spend their day. The agent connects to the contract intelligence API via MCP (Model Context Protocol), which means users can ask natural language questions like:
...and get answers without ever leaving their chat window. The agent handles authentication, queries the right endpoints, and formats responses for chat consumption.
This "meet users where they are" approach has been the difference between a tool that sits unused and one that actually gets adopted. The web UI is still there for power users who want to explore the knowledge graph or run complex cross-contract queries, but for day-to-day questions, chat is king.
The current system handles the complete contract intelligence workflow:

Multi-model OCR with dynamic schema generation, achieving comprehensive extraction across diverse document types.
PostgreSQL-first with pgvector for unified data management, eliminating synchronization issues while enabling hybrid search.
Both traditional RAG for fact-finding and GraphRAG for relationship-aware queries, with direct LLM fallback for complex cases.
Production AWS infrastructure with proper security, plus agent integration for chat-based access.
Across 125 documents
With Redis caching
With relationship tracking in GraphRAG
111 contracts fully modeled
This isn't just OCR—it's document intelligence that understands context, relationships, and the difference between "what's written" and "what's current."
The foundation is solid, but document intelligence is still early innings. The next frontier involves:

Understanding how terms change over time and surfacing trends
"Which clients have similar liability structures?"
Detecting implicit relationships that aren't explicitly stated
Notifying teams when contracts approach expiration or when terms conflict
The key insight driving everything forward: documents don't exist in isolation. They're part of networks of relationships, temporal sequences, and business contexts. The systems that crack this—that can reason about document hierarchies as easily as they extract text—will transform how organizations understand their own commitments.
From OCR to Intelligence: Building a Contract Analysis Platform That Actually Works