Part 4 of our guide to how to build an AI agent for production.
If you can't measure the quality of your AI system, every improvement you ship is luck. That's the whole premise of the hypothesis-driven approach in part 1: you can't iterate on what you can't evaluate.
The question is how. This post covers the LLM evaluation method that has become standard in production, when it's the right call, and the step-by-step process for building one you'd actually trust.
LLM evaluation: the three options
Human annotation is the gold standard. It's also slow, expensive, and hard to keep consistent. Great for a final benchmark. Useless for daily iteration.
Automated metrics like BLEU and ROUGE are fast and nearly meaningless for modern LLM outputs. They measure surface overlap with a reference answer, which is not what "good" means for an agent that can phrase a correct answer a hundred ways.
LLM-as-a-judge is the middle path: use another model to score the outputs of your system against criteria you define. Zheng et al. (2023) showed that strong LLM judges agree with human evaluators over 80% of the time on chatbot quality, roughly the same rate humans agree with each other.
It's not magic and it's not right for every problem. Here's how to tell.
When LLM-as-a-judge is the right tool
An LLM judge is cheaper and faster than humans, doesn't get bored or sandbag, and is worse at following complex instructions. That trade-off works in your favour in three situations.
1. The judgment isn't deeply intellectual. The criterion doesn't require deep domain expertise (regulatory, medical, legal nuance) or multi-step logical reasoning. An LLM judge can reliably tell you whether the answer was grounded in the provided context. It cannot reliably tell you whether the answer logically follows from the context. Know which question you're asking.
2. You need fast iteration. If you ship daily and need feedback in hours, the human loop is too slow. An LLM judge re-scores your whole test set every time you change a prompt.
3. You don't have an annotation function. Training annotators is real work: curriculum, exams, ongoing quality checks, people gaming the system. It's a department. If you don't have one, don't pretend. Start with a judge.
If none of these apply, you may actually need humans. That's a legitimate answer.
A practical LLM evaluation framework: building the judge
Building a judge you trust is the same work as building any model. You just concentrate the labelling effort into a small, focused team.
1. Find a real domain expert. Someone who knows what "good" looks like for this system. Not the engineer who wrote the prompt; someone closer to the business outcome. Together, write down the criteria: relevance, safety, factual accuracy, tone, whatever is load-bearing for your product.
2. Label a representative sample together. Pull a random slice of real production inputs, not curated ones, and have the expert score outputs against the criteria. You score them too.
3. Argue until you agree. Early disagreement is the most valuable signal you'll get. It means the criteria aren't well-defined yet, even between two humans who supposedly understand them. Re-label, sharpen definitions, repeat until you agree on a clean set. What comes out is your golden set: canonical good and bad examples.
4. Tune the judge against the golden set. One slice for tuning, another held out for testing. Options from cheapest to most expensive:
- Iterate on the judge prompt
- Decompose: one judge per criterion instead of one for everything
- Add agency: let the judge call tools, look at related data
- Fine-tune a smaller model
Your headline LLM evaluation metric is agreement with the golden set. If it's not high enough to act on, decompose further or send the hard cases to humans.
AI agent evaluation: scoring the steps, not just the answer
Evaluating a chatbot means scoring the final answer. Evaluating an agent is harder: the final answer can be right for the wrong reasons, or wrong because of one bad tool call six steps back.
In practice this means running the judge at two levels:
- Outcome. Did the final result meet the criteria?
- Trajectory. Was each step reasonable? Did the agent call the right tool with the right arguments? Did it stop when it should have?
Trajectory evaluation needs step-level traces, which is why evaluation and observability are built together, not separately. It also tells you which of the five agent components is responsible for a failure. And it's the reason self-checking orchestration patterns like Reflexion work: they run a judge inside the loop.
Why this beats hiring annotators, when it fits
The killer feature people miss is flexibility. When your product changes, your criteria change. Changing a prompt is a five-minute job. Retraining fifty annotators against a new definition is a quarter.
You can also stack: hard cases to humans, easy cases to the judge, and LLM-generated suggestions to speed the humans up. You don't pick one approach forever.
The bottom line
If your labelling task doesn't clearly fail the three conditions above, default to LLM-as-a-judge. If it doesn't hold up, split the work. Write prompts, run evals, look at where the judge disagrees with the golden set, iterate.
Once you have a judge you trust, no AI project is scary any more, because every change becomes measurable.
This post is part 4 of our guide to how to build an AI agent for production. The series:
- AI implementation strategy: why AI projects die in pilot
- AI agent architecture: the 5 components
- AI agent orchestration: the 5 patterns
- LLM-as-a-judge: evaluating LLM outputs (you are here)
- LLM observability: debugging agents in production
Trying to set up evals on a system that's already in production? That's the engagement we run most often. Talk to us about AI agent development.
