23 AI Prototypes in Three Months: What Separates the 3 That Shipped from the 20 That Didn't
Boris Cherny's team at Claude prototyped a terminal spinner 50 to 100 times. Eighty percent never shipped. Agent teams iterated through hundreds of versions before finding something production-worthy. This isn't failure. This is what real AI development looks like when you're honest about it.
Most articles about AI prototyping celebrate the winners. They interview the teams whose products made it to production, extract some wisdom about "moving fast" or "user-centered design," and call it a day. That's survivorship bias dressed up as strategy.
The real learning happens in the graveyard. After three months of evaluating 23 AI prototypes and watching 20 of them die, I can tell you exactly what separates production-viable work from expensive science experiments.
Table of Contents
- Why Most Prototype Evaluation Frameworks Are Theater
- The Four Gates Every Production AI Must Pass
- What the 20 Dead Prototypes Had in Common
- How the 3 Winners Were Actually Different
- A Framework You Can Actually Use Tomorrow
Why Most Prototype Evaluation Frameworks Are Theater
Walk into any innovation lab and ask how they evaluate AI prototypes. You'll hear about "demo days" and "stakeholder feedback sessions" and "alignment with strategic priorities." All theater.
The real question nobody asks: can this thing survive contact with production?
Most prototype reviews focus on whether the AI works in the demo. Does it answer questions correctly? Does the interface look good? Can it handle the happy path? These are necessary conditions, not sufficient ones.
Analogy: Evaluating AI prototypes on demo performance is like hiring a chef based on how their food photographs. Great plating doesn't tell you if they can handle a Friday night rush, manage inventory, or keep a kitchen from falling apart under pressure.
Production viability isn't about whether something works. It's about whether it keeps working when everything around it changes.
The Four Gates Every Production AI Must Pass
After watching 20 prototypes fail and 3 succeed, four gates emerged as non-negotiable. Miss any one and your prototype stays a prototype.
Gate 1: The Latency Reality Check
Your prototype responds in 3 seconds during the demo. Beautiful. What happens when 50 users hit it simultaneously? What about 500?
The terminal spinner Boris Cherny's team prototyped 100 times? That was about latency perception. They needed something that made LLM delays feel acceptable to developers who expect instant terminal feedback.
The test: Run your prototype at 10x expected load for 24 hours. If response times degrade by more than 50%, you have a problem. If they break completely, you have a prototype that will never ship.
Gate 2: The Hallucination Audit
Every LLM hallucinates. The question is whether your system can detect it, handle it, and recover from it without human intervention.
One of our dead prototypes was a customer support agent that answered questions correctly 95% of the time in testing. Great demo. In production simulation, the 5% hallucination rate meant one wrong answer every 20 interactions. For a support tool, that's catastrophic.
The test: Run 1,000 production-like queries through your system. Log every response. Have domain experts review 100% of them. If you can't explain and mitigate every hallucination, you're not ready.
Gate 3: The Integration Tax
AI prototypes love to live in isolation. Clean APIs. Perfect data. No legacy systems. Production is the opposite.
One prototype died because integrating it with the existing CRM required modifying 17 database tables and rewriting 3 years of business logic. The AI worked fine. The integration cost was impossible to justify.
The test: Map every system your AI needs to touch. Document every API call, database query, and data transformation. If the integration work exceeds the prototype build time by more than 2x, reconsider.
Gate 4: The Maintenance Math
Who fixes this when it breaks? Who retrains the model when performance drifts? Who handles the weekly prompt updates when the underlying LLM changes?
The question isn't whether your prototype needs maintenance. It's whether the maintenance cost is sustainable.
The test: Estimate monthly maintenance hours. Multiply by your team's hourly cost. Compare to the value delivered. If maintenance costs exceed 30% of delivered value, the economics don't work.
What the 20 Dead Prototypes Had in Common
The failures shared patterns. Here's what killed them:
Pattern 1: Demo Magic That Collapsed Under Load
Seven prototypes died because they only worked with clean, curated data. The moment real production data hit them, accuracy dropped 40% or more.
One document analysis tool worked perfectly on PDFs created in the last two years. Then someone fed it a scanned contract from 2003. The OCR failed, the extraction failed, everything cascaded.
Pattern 2: The "We'll Fix It Later" Technical Debt
Five prototypes accumulated so much technical debt during the build that shipping them would have meant accepting known bugs, security issues, or architectural problems.
One team built a working prototype in two weeks using a library they knew was deprecated. "We'll refactor later." Later never came. The library stopped being maintained. The prototype died.
Pattern 3: The Invisible Integration Iceberg
Four prototypes died when integration revealed complexity nobody saw coming. Authentication systems that needed custom logic. Data formats that required translation layers. Legacy APIs that broke under new load patterns.
The work to integrate exceeded the work to build by 5x or more.
Pattern 4: Economics That Don't Close
Four prototypes had usage costs that made them economically unviable. One prototype cost $12 in API calls to deliver $15 in value. The margin was too thin to justify maintaining it.
How the 3 Winners Were Actually Different
The three prototypes that shipped weren't necessarily better at AI tasks. They were better at surviving the four gates.
Winner 1: The Boring Integration
A document classifier that integrated with existing systems instead of trying to replace them. It didn't do anything impressive. It just did one thing reliably, with clear failure modes, in a way that fit into current workflows.
No heroic AI. No ambitious vision. Just a tool that solved a real problem without requiring anyone to change how they work.
Winner 2: The Constrained Scope
An email drafting assistant that only worked for one specific type of email: scheduling follow-ups. Not general-purpose writing. Not creative content. Just one narrow use case, done extremely well.
The constraint made it possible to train, validate, and maintain. The narrow scope meant hallucinations were easy to catch and fix.
Winner 3: The Honest Hybrid
A research assistant that knew when to use AI and when to fall back to traditional search. It didn't try to make AI do everything. When confidence was low, it showed traditional results instead of hallucinating.
Users trusted it because it was honest about its limitations.
A Framework You Can Actually Use Tomorrow
Here's the evaluation scorecard we used. Score each category out of 100. Anything below 70 overall doesn't ship.
| Category | Weight | Passing Score |
|---|---|---|
| Latency at 10x load | 25% | 75/100 |
| Hallucination rate & handling | 25% | 80/100 |
| Integration complexity | 20% | 70/100 |
| Maintenance cost ratio | 15% | 60/100 |
| Economic viability | 15% | 70/100 |
How to score each category:
Latency (25%): Run at 10x expected load for 24 hours. Score = 100 minus percentage degradation. If response time doubles, score is 50.
Hallucination (25%): Run 1,000 production queries. Score = (100 - hallucination_rate * 10) * detection_quality_multiplier. If you can't detect hallucinations, multiply by 0.5.
Integration (20%): Count integration touchpoints. Score = 100 - (number_of_systems * 5) - (data_transformations * 3). Cap at 0.
Maintenance (15%): Calculate monthly maintenance hours * hourly cost / monthly value delivered. Score = 100 - (ratio * 200). If maintenance exceeds 50% of value, score is 0.
Economics (15%): Calculate cost per transaction / value per transaction. Score = 100 - (ratio * 100). If costs exceed value, score is 0.
The Uncomfortable Truth About Prototype Velocity
Building 23 prototypes in three months sounds impressive. It's not. The impressive part is killing 20 of them quickly.
Most organizations fall in love with their prototypes. They see the demo working and assume production is just deployment. It's not. Production is where the hard work starts.
The teams that ship are the ones willing to kill prototypes fast when they fail the gates. No extended pilots. No "let's see if we can make it work." Just honest evaluation and quick decisions.
Boris Cherny's team didn't prototype a spinner 100 times because they couldn't get it right. They did it because they were willing to kill 99 versions that didn't meet the bar. That's the skill.
What This Means for Your AI Practice
If you're running an AI innovation practice, your job isn't to maximize prototypes built. It's to maximize prototypes killed before they waste resources.
Set clear gates. Apply them ruthlessly. Celebrate the kills as much as the ships. The prototype that dies in week two saves more money than the one that limps into production and fails.
The math is simple. Three months. Twenty-three prototypes. Three shipped. That's a 13% success rate. For AI prototyping, that's not bad. That's realistic.
If your success rate is higher, you're probably not pushing hard enough. If it's lower, you're probably not killing fast enough. The sweet spot is somewhere between 10% and 20%, depending on how experimental you want to be.
Conclusion
The difference between prototypes that ship and prototypes that die isn't usually about the AI. It's about latency, hallucinations, integration, maintenance, and economics.
Build your gates. Apply them early. Kill fast. Ship the 13% that survive.
That's how you build an AI practice that delivers value instead of expensive demos that never see production.