In the previous parts of this series, we built a robust enterprise infrastructure:
- Part 1: Distributed ingestion using Kafka, Redis Streams, and Gemini 3.8 Flash.
- Part 2: Canonical Common Data Models using PEPPOL and HL7 FHIR.
- Part 3: Medallion Lakehouse persistence with Apache Iceberg and Temporal DAGs.
- Part 4: Hybrid semantic search and Knowledge Graph policy verification.
Now our documents are ingested, validated, stored, and contextualized. But as documents move through the pipeline, our system must make hundreds of micro-decisions every minute:
In Accounts Payable:
- Does this invoice match the purchase order within the allowable 2% shipping variance?
- Has this vendor’s remittance bank account changed compared to the last six months (potential invoice fraud)?
- Is this invoice a duplicate submission of a bill received last week?
In Healthcare RCM:
- What is the probability that this CMS-1500 claim will be denied by UnitedHealthcare?
- Does this diagnostic code require an attached medical record before electronic submission?
- Is this claim clean enough for automated auto-adjudication, or does it require manual coding review?
If you route every single one of these micro-decisions to a massive frontier LLM like Claude 3.7 Sonnet or GPT-5, your system will be slow and financially unviable.
A single turn on a frontier model takes between 1,200ms and 2,500ms, costing roughly $3.00 per million tokens. When processing 50,000 documents a day, calling a generative LLM for every minor branch in your business logic will add over 20 hours of daily compute latency and inflate your monthly cloud bill by tens of thousands of dollars.
To build an efficient enterprise platform, you must adopt the Dual-Process Architecture: deploying a fast, dedicated System-1 Decision Model as the gatekeeper.
In this fifth installment of our AI System Design Series, we integrate TypeSafe AI’s Jev:
- The Dual-Process Computing Pattern: Applying Daniel Kahneman’s cognitive framework to distributed software.
- Architecture of Jev: Why discriminative decision models outperform autoregressive chat models on triage.
- Domain 1 (Accounts Payable): Sub-120ms 3-Way Matching and fraud anomaly screening.
- Domain 2 (Healthcare RCM): Real-time pre-submission denial risk prediction and prior-authorization triage.
- Production Implementation: Building calibrated threshold routers with
typesafe_sdk.
The Dual-Process Computing Pattern
In cognitive science, Daniel Kahneman described two modes of thought:
- System 1 operates automatically and quickly, with little or no effort and no sense of voluntary control (reflexes, immediate pattern recognition).
- System 2 allocates attention to effortful mental operations, including complex computations and deliberative logic.
In modern software engineering, generative LLMs are System 2: they are brilliant at synthesizing complex legal contracts, writing clinical appeal letters, and refactoring code, but they are heavy, slow, and expensive.
Jev is System 1: a specialized model trained specifically to evaluate an input state against an explicit set of typed questions (booleans, enums, confidence bounds) and return a calibrated decision in under 150 milliseconds.
Incoming Document State
|
v
+------------------------------------+
| System-1 Decision Gate: Jev |
| - Latency: < 120ms |
| - Cost: $0.05 / 1M tokens |
| - Output: Typed Schema + Probs |
+-----------------+------------------+
|
+---------+---------+
| |
v v
[Clean / Low Risk] [Complex Anomaly]
| |
v v
Auto-Approve to Escalate to System-2
ERP Ledger (PydanticAI / Gemini)
(85% of Traffic) (15% of Traffic)
By placing Jev as the gatekeeper at every decision junction, between 80% and 90% of routine transactions are resolved instantly for fractions of a cent. The heavy System-2 models are invoked only for the genuine 10% to 15% of edge cases that require deep reasoning.
Domain 1: Accounts Payable 3-Way Matching
In corporate procurement, the standard internal control procedure is the 3-Way Match:
- The Vendor Invoice (what the supplier claims you owe).
- The Purchase Order (what your procurement department agreed to pay).
- The Receiving Report / Goods Receipt (what the warehouse dock actually received).
If all three documents align, the invoice should be approved and scheduled for payment automatically. If there is a price discrepancy, quantity mismatch, or suspicious alteration, it must be flagged.
Here is the typed decision schema we provide to Jev:
from pydantic import BaseModel, Field
from enum import Enum
class MatchVerdict(str, Enum):
PERFECT_MATCH = "perfect_match"
WITHIN_TOLERANCE = "within_tolerance"
PRICE_MISMATCH = "price_mismatch"
QUANTITY_MISMATCH = "quantity_mismatch"
SUSPECTED_FRAUD = "suspected_fraud"
class Invoice3WayMatchDecision(BaseModel):
verdict: MatchVerdict
price_variance_percent: float = Field(..., description="Percentage variance between PO and Invoice")
bank_account_verified: bool
is_duplicate_submission: bool
confidence_score: float = Field(..., ge=0.0, le=1.0)
requires_human_approval: bool
routing_reason: str
Executing the 3-Way Match with Jev
Using typesafe_sdk, we evaluate the combined state of the three records in a single sub-120ms network turn:
import os
from typesafe_sdk import TypeSafeClient
typesafe_client = TypeSafeClient(
api_key=os.environ["TYPESAFE_API_KEY"],
model="typesafe/jev-latest"
)
def evaluate_invoice_3way_match(
invoice: dict,
purchase_order: dict,
receiving_report: dict
) -> Invoice3WayMatchDecision:
# Construct a concise state payload combining the three sources
state_payload = {
"invoice": {
"vendor_name": invoice["supplier_legal_name"],
"vendor_iban": invoice["supplier_tax_identifier"],
"total_amount": str(invoice["payable_amount"]),
"line_items": [
{"name": l["item_name"], "qty": str(l["invoiced_quantity"]), "price": str(l["price_amount"])}
for l in invoice["line_items"]
]
},
"purchase_order": {
"po_number": purchase_order.get("po_number"),
"approved_amount": str(purchase_order.get("total_amount")),
"authorized_vendor_iban": purchase_order.get("authorized_vendor_iban")
},
"receiving_report": {
"received_quantities": receiving_report.get("items_received", [])
}
}
# Execute System-1 decision evaluation
decision = typesafe_client.decide(
state=state_payload,
schema=Invoice3WayMatchDecision,
temperature=0.0
)
return decision
When an invoice arrives:
- If
verdict == MatchVerdict.PERFECT_MATCHandconfidence_score > 0.95, the invoice bypasses all human touchpoints and is posted directly to SAP. - If
verdict == MatchVerdict.SUSPECTED_FRAUD(e.g. the invoice IBAN does not match the master vendor agreement), the invoice is locked immediately, and a high-priority security event is dispatched to the audit team.
Domain 2: Healthcare RCM Claim Denial Risk Scoring
In healthcare revenue cycle management, submitting a defective claim to an insurance clearinghouse results in costly denial management cycles:
- An initial denial takes an average of 45 to 60 days to resolve.
- The administrative cost of manually reworking a denied claim averages $25 to $118 per claim.
Before any claim leaves our system as an electronic EDI 837 transaction, Jev evaluates the claim against historical denial patterns and payer-specific submission criteria.
The Healthcare Claim Triage Schema
from pydantic import BaseModel, Field
from enum import Enum
from typing import List
class DenialCategory(str, Enum):
CLEAN_CLAIM = "clean_claim"
TECHNICAL_DEFECT = "technical_defect" # Missing NPI, invalid zip, modifier syntax
MEDICAL_NECESSITY = "medical_necessity" # Diagnosis does not justify procedure
PRIOR_AUTH_MISSING = "prior_auth_missing" # Required authorization code omitted
DUPLICATE_CLAIM = "duplicate_claim"
class ClaimPreSubmissionEvaluation(BaseModel):
denial_risk_probability: float = Field(..., ge=0.0, le=1.0)
predicted_category: DenialCategory
is_submission_ready: bool
missing_required_attachments: List[str] = Field(default_factory=list)
confidence_score: float = Field(..., ge=0.0, le=1.0)
triage_action: str = Field(description="'AUTO_SUBMIT', 'AUTONOMOUS_FIX', or 'MANUAL_CODER_REVIEW'")
Triage Logic and Routing Branches
When Jev evaluates an extracted claim:
def triage_healthcare_claim(claim_data: dict, payer_profile: dict) -> ClaimPreSubmissionEvaluation:
eval_state = {
"claim": claim_data,
"payer_historical_rules": payer_profile
}
evaluation = typesafe_client.decide(
state=eval_state,
schema=ClaimPreSubmissionEvaluation,
temperature=0.0
)
return evaluation
We now have three deterministic routing branches:
Branch 1: Clean Claim (denial_risk_probability < 0.05).
The claim is immediately converted to an ASC X12 EDI 837P/837I transmission and dispatched to the clearinghouse. Over 75% of clean routine claims proceed through this path in under 200 milliseconds.
Branch 2: Technical Defect (predicted_category == DenialCategory.TECHNICAL_DEFECT).
The claim is routed to our autonomous agent layer (Part 6) to programmatically resolve the syntax issue—such as pulling a missing provider NPI from the local directory or reformatting a date—without burdening human staff.
Branch 3: Medical Necessity or Prior-Auth Missing (predicted_category == DenialCategory.MEDICAL_NECESSITY).
The claim is routed to our Hybrid Knowledge Graph RAG layer (Part 4) to retrieve clinical coverage evidence and generate a formal appeal packet before submission.
Production Performance and Economics
Let us evaluate the empirical telemetry of introducing Jev as a System-1 gatekeeper compared to a traditional monolithic LLM pipeline processing 1,000,000 documents per month:
| Metric | Monolithic Generative Architecture | Dual-Process Stack (Jev + Generative Cascade) |
|---|---|---|
| Average Gatekeeper Latency | 1,850 ms | 115 ms (16x Faster) |
| Input Token Cost per Gate | $0.0045 | $0.00008 (56x Cheaper) |
| P99 Queue Dwell Time | 14.2 seconds | 0.8 seconds |
| Monthly Routing Compute Bill | ~$4,500 | ~$120 (97% Reduction) |
| JSON Schema Parsing Regressions | 1.8% of turns | 0.0% (Deterministic Schema Guarantee) |
Summary and What Comes Next
In this fifth installment, we solved the latency and cost bottlenecks of enterprise decision-making:
- Applied the Dual-Process cognitive architecture to distributed software systems.
- Implemented TypeSafe AI’s Jev for sub-120ms, deterministic classification and triage.
- Built an automated 3-Way Matching engine for Accounts Payable with instant fraud anomaly detection.
- Deployed a pre-submission denial risk scoring gatekeeper for Healthcare RCM claims.
Now our documents are ingested, validated, stored, cross-referenced against policies, and triaged. But what happens when an invoice has an unmapped General Ledger code, or a healthcare claim requires an automated appeal letter with clinical evidence?
In Part 6 of this series, we will build the Autonomous Agent Execution Layer with PydanticAI: orchestrating type-safe multi-turn agent loops with strict tool-calling boundaries, ERP/EHR mutations, and self-healing retries.
Get Weekly AI Architect Cost & Strategy Updates
Join 14,000+ developers receiving weekly, data-driven cost-reduction blueprints and production-ready agent guidelines.