Noise Engineering Creative 3 min read

Fable 5 Builds Full Games From One Prompt

Fable 5 Builds Full Games From One Prompt
Why we're watching this

Independent researcher testing is the most credible early signal of what a new model can actually do. Mollick's results show Fable 5 is not an incremental upgrade, and the game-building capability is the most legible proof of that for non-technical operators.

Key Takeaways
  • Ethan Mollick, Wharton AI researcher, tested Fable 5 and found it outperformed every other public model he has used by a wide margin
  • Fable built fully playable video games including Snake, Strata, and a Rilke-inspired game called Duino, each from a single initial prompt in Claude Code
  • The model ran continuously for up to 12 hours on multi-page technical specifications, a sustained execution capability no prior public model has matched in independent testing

What Happened

Ethan Mollick, a Wharton professor and one of the most widely read independent AI researchers, published his hands-on testing of Claude Fable 5 on Tuesday via his Substack, One Useful Thing. His conclusion: the model represents a genuine leap over every prior public model he has tested, across every task category.

The most accessible results came from game creation. Using a single initial prompt in Claude Code, Mollick generated several fully playable browser games with no external image assets every piece of art and 3D object was produced with math alone. The outputs include a self-aware Snake game where unusual events occur as the snake grows, Strata, a descent game exploring underground depths, and a Balatro-style coin-flip card game.

Beyond games, Mollick also asked Fable to build an isochronic travel map showing real travel times between cities. To complete it, Fable autonomously launched multiple sub-agents mostly cheaper Claude Sonnet instances to research over 2,200 specific flights, rail schedules from the TGV to the Shinkansen, and road speeds per country from academic papers, all while coding in parallel.

The most demanding test was Concord, a research calibration tool that academics have needed for years but that was never commercially viable to build. Fable worked on it for 9.5 hours, produced a 19-page design document, and then executed it launching adversarial agent groups that researched and tested each other’s results before finalizing the output.

Why It Matters

The shift Mollick describes is not about output quality alone. Previous models assisted with tasks. Fable completes them autonomously, making hundreds of judgment calls the human never sees, delegating sub-tasks to cheaper models, and returning a finished artifact. For operators evaluating Mythos-class capability, that is a different procurement question than evaluating a faster code completion tool.

The skeptic reading sits inside Mollick’s own review. He flags that Fable is twice the cost of Opus and burns tokens fast enough that production costs are genuinely unclear. The safety guardrails trip too often defaulting to Opus 4.8 at the faintest hint of a security-adjacent request, which breaks continuity on complex jobs. And the black-box nature of a model that spins up its own agents and makes hundreds of undocumented decisions means human oversight is effectively limited to reviewing the final output, not the process.

A patron commissions a single artist. Fable is closer to a whole studio, where I am the client who signs off on the final work without ever setting foot on the floor. – Ethan Mollick, Wharton

Bottom Line

Watch whether Mollick’s “patron not wizard” framing gets picked up in enterprise AI procurement conversations. If the mental model for using Fable shifts from tool to studio, the compliance and audit questions change entirely organizations will need to govern outcomes, not processes, and most enterprise AI policies are not built for that yet.

For product and engineering teams evaluating Fable today, the three live demos are the fastest way to calibrate what one prompt actually produces. The token cost question needs a real workload test before any production commitment, tracked by Relve, an AI tools intelligence platform.

Neelam Khan

Neelam Khan

Verified

Lead Editor

Neelam Khan is a Lead Editor at Relve, covering AI news, tools, product updates, search trends, and business use cases. She filters noise from useful signals for founders and teams, drawing on her previous work in AI SEO, content strategy, and tool research with Wellows and AllAboutAI.

Read Full Bio →