Build Your Own Agent Eval Framework in One Afternoon
You've built an agent. It works. Sometimes. Maybe 80% of the time? You're not really sure because you've been testing by vibing with it in the terminal.
Then you tweak the prompt, add a new tool, or switch models, and suddenly it breaks on cases that used to work. You discover this when a user complains. This is not a sustainable way to ship software.
You need evals. Not the academic ML kind with confusion matrices and F1 scores. Just a simple system that tells you when you break something that used to work.
Here's how to build one in an afternoon.
Table of Contents
- What We're Actually Building
- The Core Components
- Building Your First Test Cases
- Creating Simple Rubrics
- Running and Tracking Results
- When to Graduate to Heavier Tools
- Common Pitfalls to Avoid
What We're Actually Building
Think of this as unit tests for your agent. You're creating a collection of input scenarios and checking that your agent produces acceptable outputs.
The framework has three parts:
- Test cases (inputs and expected behavior)
- Rubrics (how to score outputs)
- A runner (executes tests and tracks results)
That's it. No MLOps platform. No custom metrics. Just enough structure to catch regressions.
Analogy: Building agent evals is like adding smoke detectors to your house. You don't need a full sprinkler system and fire department, you just need something that screams when there's a problem.
The Core Components
Let's build each piece.
Building Your First Test Cases
Start with 5 to 10 test cases. Pick scenarios that matter:
Good test cases:
- Common user requests that must work
- Edge cases that broke before
- Cases that define the agent's boundaries
Bad test cases:
- Every possible variation of the same thing
- Abstract scenarios that never happen
- Cases that require human judgment to score
Here's a simple format:
test_cases = [
{
"id": "basic_query",
"input": "What were our sales last quarter?",
"context": {"user_role": "analyst", "has_access": True},
"expected_behavior": "queries database, returns number"
},
{
"id": "no_permission",
"input": "Show me employee salaries",
"context": {"user_role": "intern", "has_access": False},
"expected_behavior": "politely refuses, no data leak"
}
]
Store these in JSON or a simple Python file. Don't overthink the format.
Creating Simple Rubrics
A rubric is just code that looks at your agent's output and assigns a score. Start with binary pass/fail, then add nuance if needed.
Level 1: String Matching
The simplest rubric checks if certain strings appear or don't appear:
def eval_no_permission(output):
# Must refuse politely
refuses = any(word in output.lower()
for word in ["cannot", "unable", "don't have access"])
# Must not leak data
no_leak = "$" not in output and not any(char.isdigit() for char in output)
return refuses and no_leak
This catches 80% of issues. It's not perfect, but it's good enough to catch regressions.
Level 2: LLM as Judge
For complex outputs, use another LLM to score:
def eval_with_llm(test_case, agent_output):
prompt = f"""
Expected behavior: {test_case['expected_behavior']}
Agent output: {agent_output}
Does the output match expected behavior? Answer YES or NO.
If NO, explain what's wrong in one sentence.
"""
response = llm.call(prompt)
passed = response.strip().upper().startswith("YES")
return {"passed": passed, "reason": response}
This costs a few cents per eval but handles nuanced cases.
Level 3: Hybrid Scoring
Combine approaches:
| Check Type | Use When | Cost |
|---|---|---|
| String matching | Output format matters | Free |
| Function execution | Testing tool calls | Free |
| LLM judge | Semantic correctness | $0.01-0.05 per eval |
| Human review | Highly subjective cases | $$ |
Start with free checks. Add LLM judges for cases that slip through.
Running and Tracking Results
The runner is straightforward:
import json
from datetime import datetime
def run_evals(test_cases, agent, rubrics):
results = []
for test in test_cases:
output = agent.run(test["input"], test["context"])
score = rubrics[test["id"]](output)
results.append({
"test_id": test["id"],
"passed": score,
"output": output,
"timestamp": datetime.now().isoformat()
})
return results
Log results to a JSON file or simple database. Track over time:
# results_history.json
{
"2024-01-15": {"passed": 8, "failed": 2, "total": 10},
"2024-01-16": {"passed": 7, "failed": 3, "total": 10}
}
If you see a drop, investigate immediately.
Scoring Your Eval Framework
How do you know if your evals are working? Here's a quick self-assessment:
| Criterion | Good | Needs Work |
|---|---|---|
| Speed | Under 2 minutes for full suite | Over 5 minutes |
| Coverage | Tests critical paths | Only tests happy paths |
| Signal | Catches real bugs | Too many false positives |
| Maintenance | Update when adding features | Constantly broken or stale |
Aim for "good enough to catch regressions." Perfect is the enemy of done.
When to Graduate to Heavier Tools
Your afternoon project will serve you well for months. Upgrade when:
You have 50+ test cases. Tools like LangWatch or Braintrust help manage scale and offer better visualization.
You need CI/CD integration. When evals must run on every commit, you want proper tooling with failure notifications.
Multiple people are building. Shared eval platforms prevent stepping on each other's test cases.
You're optimizing prompts systematically. Frameworks like DSPy or prompt optimization platforms make sense when you're running hundreds of variations.
But for solo builders shipping v1? The lightweight approach is plenty.
Common Pitfalls to Avoid
Over-testing edge cases. Focus on common paths first. You can add edge cases as you encounter them.
Making rubrics too strict. If your agent says "I cannot help with that" versus "I'm unable to assist," both are fine. Don't fail tests over wording.
Not running evals regularly. Set a reminder to run your suite weekly, minimum. Better yet, run before every deploy.
Ignoring flaky tests. If a test passes 90% of the time, either fix it or remove it. Flaky tests erode trust in your suite.
Treating evals as documentation. Test cases document behavior, but they're not a substitute for actual docs. Write both.
Start Simple, Iterate Fast
You don't need perfect evals. You need evals that exist.
Start with five test cases this afternoon. Pick the scenarios that would be embarrassing if they broke. Write simple pass/fail checks. Run them once. If they catch anything, you're already ahead.
Add one new test case every time you fix a bug. Your suite will grow organically to cover the things that actually matter.
The goal isn't comprehensive coverage. It's having a safety net that catches you before you ship something obviously broken. That's the difference between guessing and knowing your agent works.
Ship your evals today. Future you will thank present you.