Theme: The Responsibility
Length: ~800 words
Hook: Analytical — expose the gap between testing and agent evaluation
Software testing is solved. We know how to test software. Unit tests, integration tests, end-to-end tests, property-based tests, mutation testing. We have frameworks, best practices, and decades of experience.
Agent evaluation is not solved. And it's not the same problem.
The fundamental difference.
Software testing asks: "Does the code produce the correct output for this input?"
The answer is binary: yes or no. The function returns 4 for input 2+2, or it doesn't. The test passes or fails.
Agent evaluation asks: "Is the agent's output correct, safe, appropriate, and useful for this input?"
The answer is not binary. It's multi-dimensional:
- Correct: Is the information accurate?
- Safe: Does it violate any policies?
- Appropriate: Is the tone, style, and content suitable for the context?
- Useful: Does it actually help the user?
- Normal operation (the happy path)
- Edge cases (unusual but valid inputs)
- Failure modes (invalid inputs, missing data, API failures)
- Adversarial inputs (prompt injection, jailbreak attempts)
- Policy violations (inputs that might trigger unsafe behavior)
An agent's response can be correct but not useful. Safe but not appropriate. Appropriate but not correct. "It works" doesn't capture any of these dimensions.
Why traditional testing doesn't work for agents.
Traditional tests assume determinism: same input, same output. Agents are non-deterministic. The same input can produce different outputs on different runs — and both might be correct.
If you write a traditional test that checks "does the agent return X for input Y," it will fail 30% of the time — not because the agent is wrong, but because the agent found a different valid path to the answer.
This means you can't use pass/fail testing for agents. You need scoring.
What agent evaluation actually looks like.
A proper agent evaluation framework has four layers:
Layer 1: Rule-based checks. Deterministic checks that don't depend on the agent's output being unique. Did the agent stay within its tool scope? Did it respect the cost ceiling? Did it complete within the time limit? Did it call only allowed APIs? These are pass/fail — and they should be enforced by policy, not by evaluation.
Layer 2: Expected-outcome matching. For tasks with known correct answers, compare the agent's output to the expected output. But use fuzzy matching, not exact matching. "The refund was processed for $49.99" and "I've issued a refund of $49.99 to your account" are both correct. Exact string matching would fail one of them.
Layer 3: LLM-as-judge. For tasks without known correct answers, use an LLM to evaluate the agent's output. The judge LLM scores the output on dimensions: accuracy, completeness, safety, appropriateness. This is probabilistic, not deterministic — but it's better than no evaluation.
The key insight: LLM-as-judge should evaluate quality, not correctness. Correctness is a rule check. Quality is a judgment. Don't ask the judge "is this correct?" — ask "is this helpful, safe, and well-reasoned?"
Layer 4: Human review. For high-stakes agents, a human reviews a sample of outputs. Not every output — a statistically significant sample. The human catches things that rules and LLMs miss: tone, cultural sensitivity, contextual appropriateness.
The evaluation suite.
Every agent in production should have an evaluation suite: 50-200 test cases covering:
The suite runs on every agent version change. The scores are tracked over time. If the score drops, the agent doesn't ship.
Why most teams don't have this.
Building an evaluation framework is hard. It's not glamorous. It doesn't produce a demo that impresses investors. It's infrastructure — the kind of thing that only matters when something goes wrong.
But when something goes wrong — and it will — the evaluation framework is the difference between "we caught this in testing" and "we shipped a broken agent to production."
The practical takeaway.
If you're building agents and you don't have an evaluation suite, you're not testing. You're hoping. And hoping is not a strategy.
Start small. 10 test cases. Rule-based checks first. Add expected-outcome matching second. Add LLM-as-judge third. Add human review for high-stakes agents.
The goal isn't perfection. The goal is to know — before production — whether your agent is getting better or worse. Without evaluation, you don't know. And what you don't know can hurt your users, your business, and your reputation.
What does your agent evaluation process look like — and is it actually catching problems before production?