Real-world enterprise documents are messy: thermal paper receipts with faded ink, crumpled delivery notes, skewed smartphone scans, and multi-column tables with merged header cells.
Here is how to design a Zero-Failure Document Extraction Pipeline that handles handwriting and complex tables without hallucinations.
1. Schema with Self-Validating Mathematical Rules
from pydantic import BaseModel, Field, model_validator
class TableRow(BaseModel):
item_code: str
description: str
quantity: float
unit_price: float
total: float
class FinancialDocument(BaseModel):
document_id: str
rows: list[TableRow]
subtotal: float
tax: float
calculated_grand_total: float
@model_validator(mode="after")
def verify_math_consistency(self):
computed_subtotal = sum(r.total for r in self.rows)
# Tolerate 1-cent rounding discrepancies
if abs(computed_subtotal - self.subtotal) > 0.05:
raise ValueError(f"Math check failed: Rows sum {computed_subtotal} != subtotal {self.subtotal}")
return self
2. Multimodal Extraction Engine
import os
from google import genai
from google.genai import types
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
def extract_handwritten_table(image_bytes: bytes) -> FinancialDocument:
prompt = (
"Transcribe this handwritten inventory dispatch document into structured JSON. "
"Calculate arithmetic totals verified against row values. If ink is smudged, infer the number mathematically."
)
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[types.Part.from_bytes(data=image_bytes, mime_type="image/jpeg"), prompt],
config=types.GenerateContentConfig(
response_mime_type="application/json",
response_schema=FinancialDocument,
temperature=0.0
)
)
return FinancialDocument.model_validate_json(response.text)
FREE CODE TEMPLATE
Download the Complete PydanticAI Document Parser Blueprint
Get the complete, type-safe invoice and ID card parsing codebase in Python + a ready-to-run Docker environment. 100% free.