Essays - Benedict Evans · Benedict Evans
articleinsightThe Deep Research problem
18 February 2025Strategy
Source excerpt
About Essays - Benedict EvansOpenAI’s Deep Research is built for me, and I can’t use it. It’s another amazing demo, until it breaks. But it breaks in really interesting ways.
This is a limited feed-provided excerpt, not the full original work.
Apollo layer
What a founder can learn
Founder takeaway
Evaluate AI products by how they fail in real workflows, not by how impressive their best demonstrations look. A compelling capability is not yet dependable customer value if users cannot recognize, correct, or recover from its errors.
Why it matters
Reliability and recoverability often determine whether customers adopt an AI feature for consequential work. Studying failure modes can reveal the product controls, transparency, and scope needed to turn a demo into a trusted tool.
Relevant guides
Put it to work
Choose one critical AI workflow in your product and document: What can go wrong, how will the user detect it, and what is the fastest path to verify or recover?