If you are testing your AI agent by opening a terminal, typing three prompts, and saying โLooks pretty good to me,โ you do not have an engineering pipelineโyou have a casino.
The moment you tweak your system prompt or swap the underlying foundation model, you have zero visibility into whether you improved reasoning or silently degraded tool execution in 15% of edge cases.
In autonomous systems, Evals are the unit tests of the AI era.
Here is how production teams build rigorous, deterministic evaluation frameworks for multi-turn agent trajectories.
1. The Anatomy of an Agent Trajectory
An agent is evaluated across two distinct dimensions:
- The Final Outcome (State Verification): Did the database record update? Was the invoice emailed?
- The Execution Trajectory (Efficiency & Safety): How many superfluous tools were called? Did it hallucinate credentials? Did it burn 40,000 tokens in an infinite recursive loop?
User Goal โโโโบ [Step 1: Plan] โโโโบ [Tool: SearchDB] โโโโบ [Step 2: Reason] โโโโบ [Tool: WriteFile] โโโโบ Goal Met
โฒ โ
โโโโโโโโโ Evaluated by Evals โโโโโโโโโโโโโ
2. The Three Tiers of Production Evals
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ THE 3-TIER EVALUATION STACK โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Tier 1: Deterministic Invariants (Fast, Free, CI/CD Gate) โ
โ - JSON Schema validity, Tool call presence, Exit code 0, Schema match โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Tier 2: State Diff Verification (Mocked Environment Inspection) โ
โ - Database mutations, Git commit diff assertions, File system checks โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Tier 3: LLM-as-a-Judge Jury (Qualitative Rubric Scoring) โ
โ - Tone, conciseness, hallucination avoidance, reasoning elegance โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Python Trajectory Test with Pytest & Mocks
import pytest
from pydantic import BaseModel
from my_agent import CustomerSupportAgent, MockDatabase
@pytest.mark.asyncio
async def test_agent_refund_policy_adherence():
# 1. Setup Mock State with known ground truth
db = MockDatabase()
db.seed_order(order_id="ORD-991", amount=45.00, status="delivered_35_days_ago")
agent = CustomerSupportAgent(db=db)
# 2. Execute Multi-Turn Task
result = await agent.run("I want an immediate refund for order ORD-991.")
# 3. Deterministic Invariant Asserts
# Rule: Orders older than 30 days CANNOT be refunded automatically
assert db.get_order("ORD-991").refunded is False, "Agent refunded an expired order!"
# Assert trajectory tool execution
trajectory = agent.get_trajectory()
tool_names = [step.tool_name for step in trajectory.tool_calls]
assert "lookup_order" in tool_names, "Agent failed to inspect order state before answering"
assert "process_refund" not in tool_names, "Agent invoked unauthorized refund tool"
assert len(trajectory.steps) <= 4, f"Agent took {len(trajectory.steps)} steps (budget: 4)"
3. How to Prevent LLM-as-a-Judge Failure Modes
When using a model (e.g., Claude 3.7 or GPT-4o) to grade another modelโs reasoning, you must guard against documented cognitive biases:
- Position Bias: Judges favor the first option presented. Solution: Run dual passes swapping output order.
- Verbosity Bias: Judges score longer, convoluted responses higher than concise answers. Solution: Enforce strict word count penalties in the rubric.
- Self-Enhancement Bias: Models rate outputs generated by their own family higher. Solution: Grade open-source models with proprietary judges and vice versa.
4. Building an Automated Eval Pipeline in GitHub Actions
Never merge a pull request to main without running a golden evaluation dataset of at least 50 representative scenarios:
- Pass Rate: Must remain $\ge 98\%$ on standard business logic.
- Step Efficiency: Average steps per task must not regress by more than 15%.
- Token Consumption: Track cost per run to catch subtle prompt bloat before it inflates cloud bills.
Get Weekly AI Architect Cost & Strategy Updates
Join 14,000+ developers receiving weekly, data-driven cost-reduction blueprints and production-ready agent guidelines.