Opus vs Sonnet: A Decision Framework for Multi-Model Agent Pipelines
You're building an agent. You need to pick a model. Claude offers Opus and Sonnet, and the internet keeps telling you Opus is "smarter" while Sonnet is "faster and cheaper." But that's not a decision framework. That's just a spec sheet.
The real question: which model makes your agent better at its job? And the answer isn't always the expensive one.
Here's how to actually decide.
Table of Contents
- The Core Trade-off: Intelligence vs Iteration
- When Sonnet with Retries Beats Opus
- The Decision Tree
- Cost Math That Actually Matters
- Real Architecture Patterns
- The Hybrid Approach
The Core Trade-off: Intelligence vs Iteration
Analogy: Choosing between Opus and Sonnet is like choosing between a senior developer who gets it right the first time and a junior developer who's fast and willing to try multiple approaches. Sometimes the senior saves time. Sometimes the junior's speed and lower hourly rate win.
Opus has better reasoning. It handles ambiguity better. It makes fewer conceptual mistakes. Sonnet is 3-4x cheaper and 2-3x faster, but it occasionally misses nuance that requires judgment.
The decision hinges on one question: does your task have a clear success criterion that you can check programmatically?
If yes, Sonnet with retries often wins. If no, you need Opus.
When Sonnet with Retries Beats Opus
Consider a code generation task where the output must pass unit tests. Sonnet might fail 30% of the time on the first attempt. But if it costs $3 per 1M tokens vs Opus at $15 per 1M tokens, you can afford five Sonnet attempts for the same price as one Opus call.
If Sonnet succeeds 70% on attempt one, 85% by attempt two, and 95% by attempt three, you're getting better economics and similar success rates.
When this pattern works:
- Output has clear pass/fail criteria
- Validation is cheap (unit tests, schema checks, API calls)
- Latency tolerance is measured in seconds, not milliseconds
- Task is well-defined with minimal ambiguity
When it doesn't:
- Judgment calls with no objective correctness
- Expensive validation (human review, production testing)
- Single-shot scenarios (customer-facing responses)
- Subtle reasoning where "close enough" isn't good enough
The Decision Tree
Here's the framework. Start at the top, follow the branches.
Step 1: Can you validate the output programmatically?
- Yes: Go to Step 2
- No: Use Opus
Step 2: What's your latency budget?
- Under 2 seconds: Use Sonnet (no retry loop)
- 2-10 seconds: Use Sonnet with retry loop (2-3 attempts)
- Over 10 seconds: Compare cost of Sonnet retries vs single Opus call
Step 3: How complex is the reasoning requirement?
- Simple generation (formatting, extraction, templates): Sonnet
- Moderate reasoning (code with context, analysis with sources): Sonnet with retries
- Deep reasoning (multi-step logic, judgment across sources): Opus
Step 4: What's the cost of a wrong answer?
- Low (internal tool, easy to fix): Sonnet
- Medium (customer-facing, reversible): Sonnet with validation
- High (financial, legal, medical): Opus
Cost Math That Actually Matters
Forget abstract benchmarks. Here's the calculation that matters for your agent.
| Scenario | Model | Attempts | Cost per Task | Success Rate |
|---|---|---|---|---|
| Simple extraction | Sonnet | 1 | $0.003 | 95% |
| Simple extraction | Opus | 1 | $0.015 | 98% |
| Code generation | Sonnet | 3 avg | $0.009 | 95% |
| Code generation | Opus | 1 | $0.015 | 97% |
| Deep analysis | Sonnet | 5 avg | $0.015 | 85% |
| Deep analysis | Opus | 1 | $0.015 | 96% |
The breakeven is usually 3-5 Sonnet attempts. If your task needs more retries than that to match Opus quality, just use Opus.
But here's the twist: many tasks don't need Opus-level quality. They need "good enough, validated, and fast."
Real Architecture Patterns
Most production agents don't use one model. They route tasks based on complexity.
Pattern 1: Retry Loop for Deterministic Tasks
for attempt in range(3):
result = sonnet.generate(task)
if validate(result):
return result
return opus.generate(task) # fallback
Use this for: code generation, data extraction, formatting, template filling.
Pattern 2: Complexity-Based Routing
if complexity_score(task) < 0.3:
return sonnet.generate(task)
elif complexity_score(task) < 0.7:
return sonnet_with_retry(task)
else:
return opus.generate(task)
Use this for: content pipelines, mixed workloads, agent orchestration.
Pattern 3: Opus for Planning, Sonnet for Execution
plan = opus.create_plan(goal)
for step in plan.steps:
result = sonnet.execute(step)
validate_and_continue(result)
Use this for: multi-step workflows, research tasks, complex automation.
The Hybrid Approach
The smartest agents don't pick one model. They use Opus where judgment matters and Sonnet everywhere else.
Use Opus for:
- Initial task decomposition
- Ambiguous requirements clarification
- Quality control on Sonnet outputs
- Customer-facing responses
- Final review before production
Use Sonnet for:
- Code generation with tests
- Data transformation
- Template-based generation
- Bulk processing
- Sub-tasks with clear specs
A typical agent might be 80% Sonnet calls, 20% Opus calls, and still maintain high quality while keeping costs reasonable.
Cost Math That Actually Matters
Let's say you're processing 10,000 tasks per day.
| Strategy | Daily Cost | Monthly Cost | Quality |
|---|---|---|---|
| All Opus | $150 | $4,500 | 97% |
| All Sonnet | $30 | $900 | 88% |
| Sonnet + 3 retries | $90 | $2,700 | 94% |
| Hybrid routing | $65 | $1,950 | 95% |
The hybrid approach: route simple tasks to Sonnet, use Opus for 15-20% of tasks that need deeper reasoning. You get 95% of Opus quality at 40% of the cost.
The Real Decision
Stop thinking about models as "smart vs cheap." Think about them as tools in a pipeline.
Sonnet is your workhorse. Fast, cheap, good enough for most tasks. Opus is your quality control. Use it where mistakes are expensive or judgment is required.
The best agents use both. They route intelligently. They validate outputs. They fail over to the stronger model when the cheaper one struggles.
Build your decision tree based on your actual tasks, your actual validation methods, and your actual cost constraints. Run experiments. Measure success rates. Adjust the routing rules.
There's no universal answer. But there is a framework. Use it.