Sakana Fugu: AI Model That Matches Fable 5 on Key Benchmarks
Sakana's Fugu isn't a bigger model — it's a conductor. One endpoint that answers easy questions itself and hands hard ones to a team of specialists. Fugu Ultra hits 73.7 on SWE-Bench Pro, beating GPT-5.5, Opus, and Gemini — but it still rents its brains.
There's a new model out of Japan — Fugu, from Sakana AI — and on the hardest coding benchmarks it beats GPT-5.5, Opus 4.8, and Gemini, while going toe-to-toe with Fable 5, the best AI on earth. The twist is that Fugu isn't really a model at all. Here's the whole story.
Why this matters now#
The two best models on the planet — Fable 5 and Mythos — just got locked behind US export controls. That leaves Fugu as the most capable model a lot of the world can actually use. So the obvious question is how something out of a smaller lab is keeping pace with the frontier.
Not a bigger model — a conductor#
The bet behind Fugu is simple: what if you didn't need a bigger brain, just a better way to coordinate the brains that already exist? Sakana's own description is blunt about it — Fugu "dynamically coordinates and orchestrates a diverse pool of powerful models." It owns no weights of its own. It conducts.
How it works#
One request, one endpoint:
- Easy questions — it answers them itself.
- Hard questions — it assembles a team: a thinker, a worker, and a verifier — and hands back one clean answer.
Under the hood it calls other models, assigns them those roles, and lets them talk to each other in plain language. The role-assignment comes from two ICLR 2026 papers Sakana published alongside it — TRINITY (the lightweight coordinator) and Conductor (trained by reinforcement learning to discover its own coordination strategies). And if one provider gets cut off, Fugu simply routes around it to another. (If "give it a goal and it finishes the job" is new to you, start with our explainer on AI agents.)
The numbers#
This is the part that turns heads. Among the models you can actually use, Fugu Ultra tops every coding and reasoning benchmark Sakana published — out-scoring Opus 4.8, GPT-5.5, and Gemini 3.1 Pro across the board.
| Benchmark | Fugu Ultra | GPT-5.5 | Opus 4.8 | Gemini 3.1 | Fable 5* |
|---|---|---|---|---|---|
| SWE-Bench Pro | 73.7 | 58.6 | 69.2 | 54.2 | 80.0 |
| TerminalBench 2.1 | 82.1 | 78.2 | 74.6 | 70.3 | 80.4 |
| LiveCodeBench | 93.2 | 85.3 | 87.8 | 88.5 | — |
| Humanity's Last Exam | 50.0 | 41.4 | 49.8 | 44.4 | 53.3 |
| GPQA-D | 95.5 | 93.6 | 92.0 | 94.3 | — |
The orchestra beats the soloists — that's the headline. The honest asterisk:
Fable 5 still edges Fugu on SWE-Bench Pro and Humanity's Last Exam (the *
column) — but it's the export-controlled model nobody can buy, which is exactly why
Fugu's "best one you can run" claim holds.
The honest catch#
One thing the benchmark flex won't tell you: Fugu doesn't own any of those brains. It rents them from the same vendors everyone else uses. So what it really buys you is resilience — reroute when a provider goes dark — not independence. And the scores above are Sakana-reported, not yet independently verified, so treat them as a claim, not a closed case.
Still, the idea holds: coordinating models can out-code any single one of them. Looks like one. The power of many.
Watch the full video
@thekernelcast on YouTube

Claude Fable 5: The Model Too Dangerous to Ship — Until Now
Anthropic just shipped Claude Fable 5 — the first public model in a new Mythos tier above Opus. It found 271 zero-days in Firefox, tops every coding benchmark, and costs 2× Opus. Here's what actually matters.
read the take→
AI Agent in 100 Seconds
AI agents explained in 100 seconds — from keyword-matching chatbots to autonomous, goal-completing systems. Same chat window, completely new engine.
read the take→
Top 5 AI Customer Support Platforms of 2026 (Ranked & Compared)
Choosing the right AI customer support platform isn't just about rankings—it's about finding the one that fits your business. In this video, we compare the top 5 AI customer support platforms of 2026, highlighting their strengths, use cases, and what makes each one stand out. Watch the full ranking to see which platform takes the #1 spot.
read the take→