Theme: The Responsibility

Length: ~700 words

Hook: Provocative — challenge the trend of using AI to govern AI


There's a growing trend in AI governance: using AI models to evaluate other AI models. An LLM checks if the agent's output is safe. An AI classifier detects if the agent's behavior is appropriate. An AI judge scores the agent's response quality.

This sounds reasonable. If AI can generate content, AI can evaluate content. If AI can make decisions, AI can review decisions.

We think this is dangerous. Here's why.

Non-deterministic governance is an oxymoron.

Governance exists to provide predictability. "If the agent does X, the governance layer will catch it." This promise requires the governance layer to behave the same way every time. If the governance layer is an AI model, it might catch X today and miss it tomorrow.

Imagine a security camera that only detects intruders 92% of the time. Would you trust it to secure your building? No — because the 8% it misses is the 8% that matters. You'd install deterministic sensors: motion detectors, door contacts, glass break sensors. These sensors trigger every time. They don't have off days.

Agent governance should work the same way. The policies that prevent an agent from accessing unauthorized data should be deterministic. The rules that enforce cost ceilings should be deterministic. The checks that require human approval for sensitive actions should be deterministic.

The appeal of AI-as-judge.

We understand why teams use AI for governance. It's flexible. You can describe what "safe" means in natural language and let the AI figure it out. You don't have to enumerate every possible policy violation. The AI can catch things you didn't think of.

This is real value — for evaluation, not for governance.

Evaluation vs. governance: the critical distinction.

Evaluation asks: "How good was this output?" This is inherently subjective. An LLM-as-judge is appropriate here because quality is a judgment, not a rule.

Governance asks: "Was this action allowed?" This is inherently objective. The action either violated a policy or it didn't. An LLM-as-judge is inappropriate here because the answer should be deterministic.

The confusion comes when teams use AI-as-judge for governance because they haven't defined their policies clearly enough. If you can't write a rule that says "the agent is not allowed to call Database.write without approval," you use an AI to "check if the agent's database access was appropriate." But "appropriate" is subjective — and subjective governance is unpredictable governance.

What deterministic governance looks like.

Policy-as-code means your governance rules are written in code, not in natural language. They're evaluated by a rules engine, not by an AI model. They produce the same result every time.

Example policies:

policy RefundApproval {
  if action.tool == "BillingAPI.refund" and action.amount > 100 {
    require_approval from: "manager"
    timeout: 3600s
    on_timeout: "deny"
  }
}

policy DataAccess {
  if action.tool matches /Database\.(write|delete)/ {
    require_scope: "data:write"
    audit: true
  }
}

policy CostCeiling {
  if agent.session_cost > 10.00 {
    deny: "Session cost ceiling exceeded"
    alert: "[email protected]"
  }
}

These policies are deterministic. They trigger every time the condition is met. They don't have false negatives. They don't have off days. They don't hallucinate.

When AI-as-judge is appropriate.

AI-as-judge has a place — in evaluation, not governance: