Noise Ops Engineering 2 min read

Sakana Fugu Ultra Beats Fable on Benchmarks

Sakana Fugu Ultra Beats Fable on Benchmarks
Why we're watching this

Sakana AI launched Fugu Ultra today and explicitly positioned it as a frontier alternative to Anthropic's government-banned models. It runs on a standard API, and as a Japanese company, Sakana sits outside the US export control directive that pulled Fable 5 and Mythos 5 from the market.

Key Takeaways
  • Fugu Ultra outscores Anthropic Opus 4.6, Gemini 3.1 high, and GPT 5.4 high on all three published benchmarks in Sakana’s own testing.
  • Fugu Ultra is not a single model. It is a multi-agent orchestration system that dynamically coordinates a pool of frontier models, built on two peer-reviewed ICLR 2026 papers.
  • OpenAI-format API compatibility means teams already using Claude, GPT, or Gemini via API can integrate it without changing their existing code structure.
  • As a Japanese company, Sakana AI is not subject to the US export control directive that forced Anthropic to pull its two newest models from the market.

What Happened

Sakana AI, a Japanese lab, launched Fugu Ultra today and explicitly positioned it as a frontier alternative to Anthropic’s government-banned models, Fable 5 and Mythos 5.

Fugu Ultra is not a single model. It is a multi-agent orchestration system that dynamically coordinates a pool of frontier models, built on two papers accepted at ICLR 2026.

Sakana’s own testing shows Fugu Ultra scores 95.1 on GPQAD, 93.2 on LCBv6, and 54.2 on SWEPro, placing it above Anthropic Opus 4.6, Gemini 3.1 high, and GPT 5.4 high on all three benchmarks.

Fugu Ultra is available via an OpenAI-format API. Teams using Claude, GPT, or Gemini can integrate it without changing their existing code structure.

Why It Matters

For SaaS teams currently building on Anthropic’s top-tier models, Fugu Ultra is a practical option worth evaluating. The published numbers put it ahead of every model it was measured against, including the full Opus 4.6, and the drop-in API means testing it costs very little.

The benchmarks come from Sakana’s own testing, not independent third-party evaluation, and the model pool Fugu coordinates is not publicly disclosed. Teams considering a switch should run it against their own use cases before treating the benchmark claims as settled.

Our Fugu Ultra model matches the performance of Fable and Mythos, delivering frontier capability without the risk of export controls. Sakana AI, @SakanaAILabs

Sakana Fugu Multi Agent Orchestration Model From Sakana AI Matches Leading AI Benchmarks
by
u/techspecsmart in
aicuriosity

Bottom Line

The export control gap Sakana is targeting is real but may not stay open. The US government could extend restrictions to non-US AI providers, and the broader export control policy is still developing. SaaS teams should not build long-term dependency on any single model’s regulatory status.

In the short term, Fugu Ultra is worth testing for Engineering teams that need Fable-level capability without access to the restricted models. The OpenAI-format API makes the comparison cheap and low-risk to run.

Neelam Khan

Neelam Khan

Verified

Lead Editor

Neelam Khan is a Lead Editor at Relve, covering AI news, tools, product updates, search trends, and business use cases. She filters noise from useful signals for founders and teams, drawing on her previous work in AI SEO, content strategy, and tool research with Wellows and AllAboutAI.

Read Full Bio →