Junior developers believe that controlling an AI model is an exercise in prompt engineeringโfinding the magical sequence of adjectives, emojis, and threats like โI will tip you $200 if you format this correctly.โ
Senior engineers know that prompt engineering is an anti-pattern.
An LLM is a non-deterministic token probability distribution. Attempting to build mission-critical enterprise banking, healthcare, or logistics software on raw natural language promises is architectural malpractice.
To build software that does not break at 3 AM, you must harness the model.
1. The Anatomy of an LLM Harness
Harnessing means wrapping an untrusted statistical engine inside deterministic software boundaries:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ THE DETERMINISTIC HARNESS โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ 1. Constrained Input Specification (Validated Pydantic Schemas) โ
โ 2. Grammar-Guided Token Masking (Logit Level: Outlines / GBNF) โ
โ 3. Deterministic Runtime Invariant Assertions โ
โ 4. Type-Safe Self-Correcting Execution Loops โ
โ 5. Cryptographic Action Guardrails (Least Privilege) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
2. Zero-Tolerance Structured Outputs (Grammar-Level Masking)
Never ask an LLM for โvalid JSONโ in natural language. If the model generates a single stray backtick or unescaped quote, your JSON parser throws JSONDecodeError.
Modern inference engines (vLLM, SGLang, llama.cpp) support CFG (Context-Free Grammar) Masking:
\[P(\text{Token}_i) = 0 \quad \forall \text{ Tokens that violate JSON Grammar Schema}\]At step $i$, if the only valid characters allowed by the schema are a closing bracket } or a comma ,, the probability of every other token in the 128,000-token vocabulary is masked to $-\infty$. Hallucinating syntax becomes mathematically impossible.
Production Implementation with PydanticAI
from pydantic import BaseModel, Field, field_validator
from pydantic_ai import Agent, ModelRetry
class WireTransferOrder(BaseModel):
recipient_iban: str = Field(..., description="Standardized IBAN")
amount_cents: int = Field(..., gt=0, le=10_000_00, description="Amount in USD cents")
routing_code: str = Field(..., min_length=9, max_length=9)
@field_validator("recipient_iban")
def validate_iban(cls, v: str) -> str:
clean = v.replace(" ", "").upper()
if not clean.isalnum() or len(clean) < 15:
raise ValueError(f"Invalid IBAN structure: '{v}'")
return clean
harnessed_agent = Agent(
"google-gla:gemini-2.5-flash",
result_type=WireTransferOrder,
retries=3,
system_prompt="You are a strict financial transaction extractor."
)
3. Dynamic ModelRetry: Self-Healing Assertion Gates
When semantic logic fails (e.g., an account balance is insufficient or business invariants are breached), do not crash the service. Feed the exact validation error back to the modelโs scratchpad:
@harnessed_agent.result_validator
def validate_business_logic(ctx, result: WireTransferOrder) -> WireTransferOrder:
# Deterministic database lookup
active_balance = 5_000_00 # $5,000 in cents
if result.amount_cents > active_balance:
# Forces the LLM to inspect the constraint and adjust
raise ModelRetry(
f"Insufficient funds: Requested {result.amount_cents} cents, "
f"but customer balance is {active_balance} cents. Adjust the transfer amount."
)
return result
4. The 4 Golden Rules of LLM Harnessing
- Never Accept Free-Text When a Categorical Exists: If an answer can be one of 5 enum values, enforce an
Enumschema. Never parse strings with.lower().startswith(). - Constrain Temperature by Task Type:
- Extraction & Tool Calling:
temperature = 0.0 - Reasoning / Synthesis:
temperature = 0.2 - 0.4 - Creative Brainstorming:
temperature = 0.7 - 0.9
- Extraction & Tool Calling:
- Decouple Thinking from Tool Arguments: Use reasoning tokens (scratchpads) to allow the model to formulate logic, but bind tool invocation exclusively to structured JSON arguments.
- Enforce Hard Token Budgets: Always set
max_tokenslimits. A rogue loop generating infinite repeating sequences will deplete your rate limits and drain your budget within minutes.
Download the Complete PydanticAI Document Parser Blueprint
Get the complete, type-safe invoice and ID card parsing codebase in Python + a ready-to-run Docker environment. 100% free.