Lenny's Newsletter · Claire Vo
podcastplaybookClaude Opus 5 review: this model is brilliant (but annoying)
24 July 2026Product
Source excerpt
About Lenny's NewsletterWatch now | 🎙️ I ran Opus 5 through my seven-model benchmark, compared it with six other leading models, and came away with a verdict I genuinely didn’t expect
This is a limited feed-provided excerpt, not the full original work.
Apollo layer
What a founder can learn
Founder takeaway
Evaluate AI models with a repeatable benchmark based on your real workflows—not reputation or raw capability alone—and assess both output quality and day-to-day friction.
Why it matters
A model can be highly capable yet frustrating to use. Structured, workflow-specific testing helps founders choose tools that improve actual performance rather than merely looking impressive in demos.
Relevant guides
Put it to work
Choose one recurring, high-value task and test several models on the same inputs. Score output quality, reliability, speed, cost, and usability before deciding which one belongs in your workflow.
Private notes