AI System Design Series (Part 7): Human-in-the-Loop Orchestration and the Active Learning Flywheel

AI System Design Series (Part 7): Human-in-the-Loop Orchestration and the Active Learning Flywheel

(Updated: ) ๐Ÿ“– 5 min read

In the first six parts of this series, we designed an end-to-end autonomous document processing platform: from distributed Kafka ingestion and Gemini 3.8 Flash extraction to Apache Iceberg lakehouses, hybrid search, Jev System-1 triage, and PydanticAI agents.

On paper, this system appears capable of running an entire enterprise without human intervention.

In production reality, claiming 100% full autonomy in financial accounting or healthcare billing is an immediate red flag.

No chief financial officer will allow an AI agent to execute million-dollar wire transfers without human approval. No hospital chief medical officer will permit an AI model to unilaterally file legal appeals with Medicare without physician review. Furthermore, in the European Union, statutory regulations (such as Article 14 of the EU AI Act) explicitly mandate effective human oversight for high-risk automated decision-making.

More importantly, human reviewers are not merely a compliance safety netโ€”they are your highest-value data asset.

When an experienced certified medical coder or accounts payable controller reviews a flagged edge case and overrides an AI decision, that human correction represents high-signal ground truth. If that correction is merely saved to an isolated database table and forgotten, your AI models will continue making the exact same mistake every week.

In this seventh installment of our AI System Design Series, we build the Human-in-the-Loop (HITL) Architecture and the Active Learning Flywheel:

  1. Enterprise Review Thresholds: Establishing deterministic policy gates where autonomy must yield to human judgment.
  2. Durable Workflow Pauses: Using Temporal Signals to pause pipelines for hours or days with zero CPU consumption.
  3. The Active Learning Flywheel: Transforming human review overrides into vector-embedded exemplar memories.
  4. Dynamic Few-Shot Prompt Hydration: Injecting relevant historical human corrections into runtime agent contexts to prevent recurring errors.

Designing Enterprise Review Thresholds

Human review should never be arbitrary. We define explicit, deterministic business rules that automatically halt execution and enqueue documents into specialized reviewer workbenches:

Automated Pipeline Execution
             |
             v
+------------------------------------+
|  Deterministic Policy Gates        |
|  - Amount > $10,000?               |
|  - Unknown Vendor Bank Account?    |
|  - Jev Confidence Score < 0.85?    |
|  - Medical Modifier Conflict?      |
+-----------------+------------------+
                  |
        +---------+---------+
        |                   |
    [PASS]               [TRIGGER]
        |                   |
        v                   v
Proceed to ERP/EHR      Temporal Workflow Pauses
Auto-Adjudication       Enqueues to Reviewer Workbench

Domain 1: Accounts Payable Invariant Gates

  1. Monetary Threshold: Invoices exceeding $10,000 require department head review; invoices exceeding $100,000 require VP/CFO authorization.
  2. Vendor Bank Account Changes: If a supplier invoice lists a new IBAN or routing number not present in the master ERP vendor file, the document is immediately frozen to prevent wire diversion fraud.
  3. Pricing Tolerance: Any line-item price variance greater than 2% against the contractual purchase order triggers human buyer review.

Domain 2: Healthcare RCM Invariant Gates

  1. Clinical Medical Necessity Denials: Denials based on clinical rationale (e.g., experimental treatment clauses) require review by a licensed physician or clinical documentation specialist.
  2. High-Dollar Inpatient Claims: Hospital bills (UB-04) exceeding $25,000 require manual auditing prior to clearinghouse submission.
  3. Disputed Billing Modifiers: When an agent suggests unbundling codes using Modifier 59, a certified coder must inspect operative notes to verify distinct anatomical sites.

Durable Workflow Pauses with Temporal Signals

In traditional systems, implementing human review requires messy polling scripts or long-lived database flags that break distributed state.

With Temporal, a workflow is durable code. When a human review threshold is triggered, the workflow simply registers a signal listener and yields execution. Temporal persists the workflow state to disk and releases all compute resources:

from temporalio import workflow
from temporalio.common import RetryPolicy
from datetime import timedelta
from pydantic import BaseModel
from typing import Optional

class HumanReviewPayload(BaseModel):
    reviewer_id: str
    decision: str # "APPROVED", "REJECTED", "CORRECTED"
    corrected_data: Optional[dict] = None
    override_reason: Optional[str] = None

@workflow.defn
class ResilientInvoiceWorkflow:
    def __init__(self):
        self.review_decision: Optional[HumanReviewPayload] = None

    @workflow.signal
    def submit_human_review_signal(self, payload: dict):
        """
        Dispatched by the frontend review dashboard when a human
        accountant clicks Approve, Reject, or updates fields.
        """
        self.review_decision = HumanReviewPayload.model_validate(payload)

    @workflow.run
    async def run(self, document_id: str, tenant_id: str) -> str:
        # Step 1: Run Ingestion and Gemini 3.8 Flash Extraction
        extracted_invoice = await workflow.execute_activity(
            extract_invoice_activity, document_id, start_to_close_timeout=timedelta(minutes=2)
        )

        # Step 2: Evaluate Policy Gates
        requires_review = (
            float(extracted_invoice["payable_amount"]) > 10000.00 or
            extracted_invoice.get("bank_account_verified") is False
        )

        if requires_review:
            # Emit review task to the human reviewer queue
            await workflow.execute_activity(
                publish_review_task_activity,
                {"document_id": document_id, "data": extracted_invoice},
                start_to_close_timeout=timedelta(seconds=30)
            )

            # Workflow pauses here. Zero CPU consumed.
            # Suspended until the human dashboard sends the signal (up to 7 days).
            await workflow.wait_condition(
                lambda: self.review_decision is not None,
                timeout=timedelta(days=7)
            )

            if self.review_decision.decision == "REJECTED":
                return f"Invoice rejected by human reviewer {self.review_decision.reviewer_id}."

            if self.review_decision.decision == "CORRECTED":
                # Use human-corrected data for downstream processing
                extracted_invoice = self.review_decision.corrected_data
                
                # Capture the human correction for the Active Learning Flywheel!
                await workflow.execute_activity(
                    record_human_correction_activity,
                    {
                        "document_id": document_id,
                        "original_data": extracted_invoice,
                        "corrected_data": self.review_decision.corrected_data,
                        "reason": self.review_decision.override_reason
                    },
                    start_to_close_timeout=timedelta(minutes=1)
                )

        # Step 3: Mutate ERP / Post Transaction
        return await workflow.execute_activity(
            post_to_erp_activity, extracted_invoice, start_to_close_timeout=timedelta(minutes=2)
        )

The Active Learning Flywheel

Now let us examine the core machine learning problem: How do we prevent our AI models from making the same mistake repeatedly?

Traditional machine learning teams assume that the only way to improve model accuracy is to retrain or fine-tune weights on a GPU cluster every month. However, fine-tuning modern frontier models is expensive, slow, and risks catastrophic forgetting.

In 2026, the dominant enterprise pattern is Dynamic Few-Shot Exemplar Hydration:

Human Reviews and Overrides AI Prediction
                  |
                  v
[Sanitization & PII Scrubbing]
                  |
                  v
[Generate Semantic Vector Embedding]
                  |
                  v
[Store in pgvector corrections_exemplar_store]
                  |
                  v
Future Ingestion Run: Semantic Similarity Lookup
                  |
                  v
Hydrate System Prompt with Relevant Past Human Fix:
"Warning: For Vendor Acme Corp, lines referencing 'Cloud Support' 
must be allocated to GL Account 6020-SOFTWARE per Controller Override."

Database Schema for the Human Correction Store

We store corrections in a dedicated corrections_exemplar_store table in PostgreSQL:

CREATE TABLE corrections_exemplar_store (
    correction_id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    tenant_id VARCHAR(64) NOT NULL,
    domain VARCHAR(32) NOT NULL, -- 'invoices', 'healthcare_rcm'
    entity_key VARCHAR(128) NOT NULL, -- e.g. Vendor Tax ID or Hospital Department
    original_extracted_value TEXT NOT NULL,
    human_corrected_value TEXT NOT NULL,
    human_explanation TEXT NOT NULL,
    
    -- Context embedding of the error context (1536 dims)
    context_embedding vector(1536),
    
    created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP
);

CREATE INDEX idx_corrections_embedding 
ON corrections_exemplar_store 
USING hnsw (context_embedding vector_cosine_ops);

Implementing Dynamic Few-Shot Prompt Hydration

When an extraction or agent worker begins processing a new document, it embeds the vendor name and line item descriptions, queries the corrections_exemplar_store for the top 3 most semantically relevant historical corrections, and dynamically injects them into the system prompt:

from google import genai
from pydantic_ai import Agent

async def retrieve_relevant_human_corrections(
    tenant_id: str, 
    document_context: str
) -> List[str]:
    # 1. Embed current document context
    embedding = await generate_embedding(document_context)
    
    # 2. Query top 3 nearest historical human corrections in pgvector
    async with db_pool.acquire() as conn:
        rows = await conn.fetch("""
            SELECT human_corrected_value, human_explanation
            FROM corrections_exemplar_store
            WHERE tenant_id = $1
            ORDER BY context_embedding <=> $2::vector
            LIMIT 3;
        """, tenant_id, embedding)
        
    return [
        f"- Correction: {r['human_corrected_value']}. Reason: {r['human_explanation']}"
        for r in rows
    ]

async def build_hydrated_agent(tenant_id: str, document_context: str) -> Agent:
    corrections = await retrieve_relevant_human_corrections(tenant_id, document_context)
    
    exemplar_block = "\n".join(corrections) if corrections else "No specific past corrections."
    
    system_prompt = f"""
    You are an enterprise accounting allocation agent.
    Allocate line items to GL accounts based on company policy.
    
    CRITICAL PAST HUMAN CORRECTIONS (Do not repeat these historical errors):
    {exemplar_block}
    """
    
    return Agent("google-gla:gemini-3.8-flash", system_prompt=system_prompt)

By hydrating the runtime context with past human decisions:

  • The system self-heals within seconds of a human correcting an error.
  • Zero GPU retraining or model fine-tuning is required.
  • Human feedback immediately prevents identical errors across millions of future documents.

Summary and What Comes Next

In this seventh installment, we bridged the gap between machine autonomy and human judgment:

  • Established deterministic review thresholds across Accounts Payable and Healthcare RCM.
  • Implemented Temporal Signals to pause distributed workflows for days with zero compute overhead.
  • Engineered the Active Learning Flywheel, turning human corrections into high-value vector embeddings.
  • Deployed Dynamic Few-Shot Prompt Hydration to eliminate recurring AI hallucinations in real time.

Now our system is autonomous, resilient, and continuously learning from human experts. But in financial accounting and healthcare, regulators and auditors enforce strict legal mandates: How do you prove what happened years later during an audit? How do you guarantee multi-tenant data segregation across hospitals and corporate subsidiaries?

In Part 8 of this series, we will build Bi-Temporal Audit Trails, Multi-Tenancy, and HIPAA/SOC-2 Data Isolation: engineering Iceberg time-travel queries and cryptographic data perimeters.

WEEKLY NEWSLETTER

Get Weekly AI Architect Cost & Strategy Updates

Join 14,000+ developers receiving weekly, data-driven cost-reduction blueprints and production-ready agent guidelines.

Professor XAI
Professor XAI ML Engineer passionate about advancing AI technologies and building intelligent systems.
comments powered by Disqus