Skip to main content
株式会社オブライト
AI2026-09-126 min read

Fugu Max & Ultra v2: Sakana AI Pricing, Benchmarks (2026)

Updated Sept 2026: Sakana AI Fugu Max ($2/$6 per 1M tokens) and Fugu Ultra v2 ($5/$30) are orchestrator models routing across a model pool. Benchmarks too.


Sakana AI announced two new models, Fugu Max and Fugu Ultra v2, on September 10-11, 2026. Fugu Max prices at $2 input / $6 output per 1M tokens, while Fugu Ultra v2 prices at $5 input / $30 output. Neither is a single giant model — both are trained orchestrators that route tasks to a pool of models behind one API. Updated September 2026, here is a fast rundown of pricing, benchmarks, and when to use which.

Pricing at a glance

ItemFugu MaxFugu Ultra v2
Input (per 1M tokens)$2$5
Output (per 1M tokens)$6$30
Cache readNot disclosed$0.50/1M tokens
Web searchNot disclosed$10 per 1,000 calls
Context lengthNot disclosed1M tokens
Max outputNot disclosed~128K tokens
AvailabilityHosted API onlyHosted API only

Sakana AI claims Fugu Max's output pricing is 40-60% cheaper than Sonnet 5, Kimi K3, and GPT-5.6 Terra. This is Sakana AI's own claim, and no independent third-party verification is available yet.

What is an orchestrator model

Fugu Max and Ultra v2 depart from the conventional single-network LLM design that answers in one forward pass. Behind a single API endpoint sits a pool of open-weight and task-specialized models, and a trained orchestrator routes each request to the model it judges best suited, recursively calling itself when needed to verify and combine outputs from lower-level models before returning a final answer. This originates with the original Fugu launched in June 2026; the key change in this update is the split into a cost-focused Max and a performance-focused Ultra v2.

Fugu Max's model pool is designed for cost efficiency, routing to "the cheapest model that can solve it," and includes a wide range of candidates including the NVIDIA Nemotron family. Fugu Ultra v2, by contrast, prioritizes hard, multi-step tasks and reportedly deliberately excludes Fable 5, Fable 5.1, and GPT-6-Astra from its pool. The technical reasoning behind that exclusion has not been disclosed.

A flow diagram: a client sends one request through an OpenAI-compatible API, the Fugu orchestrator judges task difficulty and routes it across a pool of open-weight, specialized and NVIDIA Nemotron models, recursively calling itself when needed, then verifies and merges the outputs into a single response back to the client.

Fugu Max benchmarks

According to Sakana AI, Fugu Max posted the top score on six benchmarks — Terminal Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench, and SWEFish — and claims to extend the cost-performance frontier (achieving the same accuracy more cheaply) on 7 of 10 benchmarks tested.

MetricFugu Max result
Top-score benchmarks6/6 (Terminal Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench, SWEFish)
Cost-performance frontier extended7 of 10 benchmarks
Price advantage vs. top-tier modelsClaimed 2-6x cheaper

Fugu Ultra v2 benchmarks

Fugu Ultra v2 reportedly scored best or tied-best on 5 of 8 benchmarks. Notably it scored 48.3 on Chartography (a chart/diagram comprehension benchmark), well above the 27.3 comparison score for Opus 5. On the software-engineering benchmark DeepSWE it scored 74.3, and Sakana AI claims this beats models priced 3-5x higher.

MetricFugu Ultra v2 result
Best/tied-best benchmarks5/8 (GDP.pdf, Chartography, SWEFish, DeepSWE, Toolathon)
Chartography48.3 (Opus 5: 27.3)
DeepSWE74.3 (claimed to beat models 3-5x more expensive)

What changed from the previous Fugu (June 2026)

The original Fugu and Fugu Ultra, launched June 22, 2026, were a single orchestrator lineage. This update splits that into a cost-focused Max and a performance-focused Ultra v2, each with its own pricing structure, model pool composition, and exclusion policy. Details such as parameter counts or training data beyond what is listed here remain officially undisclosed.

Getting started

Fugu Max and Ultra v2 are served through an OpenAI-compatible API. If you already use the OpenAI SDK or a compatible library, the typical setup is simply to swap the base URL to Sakana AI's endpoint and set your API key. Availability through aggregators such as OpenRouter is possible but not yet confirmed — check the latest official announcements for specific access channels. Since weights are not released, local execution is not possible; usage is always through the hosted API.

Choosing between models

For broader context on managing spend, see this AI API cost optimization strategy guide. Here is how Fugu Max and Ultra v2 compare with other leading models on price and fit.

ModelInput/Output ($/1M tokens)Best fit
Fugu Max$2 / $6Cost-sensitive high-volume tasks, everyday coding and QA
Fugu Ultra v2$5 / $30Hard multi-step reasoning, long context (1M tokens) needs
Sonnet 5$3 / $15General-purpose development, balanced use
Opus 5$5 / $25Accuracy-critical, high-difficulty tasks
GPT-6-Astra$10 / $50Latest general-purpose flagship use cases

Comparison prices are the standard published API rates as of September 2026; batch APIs, caching and discounts change the effective cost, so check each vendor's official pricing page. Note that Fugu Max output at $6 is roughly 60% below Sonnet 5 output at $15, which is consistent with Sakana AI's claimed 40-60% saving.

Another orchestrator-style offering worth knowing is Sakana Marlin, a deep-research agent. It's built for a different job than the Fugu line, so it's worth checking the intended use case when deciding between them.

Caveats and limitations

- Hosted API only: weights are not released, so no local or self-hosted deployment is possible
- Not available in the EU/EEA
- Output reproducibility: since routing can select a different model per request, output style may vary even for identical prompts
- Vendor lock-in: the model pool composition and routing logic are managed entirely by Sakana AI and cannot be controlled by users
- Hard to estimate cost precisely: effective cost depends on which model a given request is routed to, making upfront estimates difficult

Which is cheaper, Fugu Max or Fugu Ultra v2?

Fugu Max is cheaper, at $2 input / $6 output per 1M tokens versus Fugu Ultra v2's $5 input / $30 output. Ultra v2 is positioned as stronger for hard tasks and offers a 1M-token context window.

Can Fugu Max or Ultra v2 run locally?

No. Weights are not released, and the models are only accessible through the hosted API.

Are Fugu Max and Ultra v2 available outside the EU/EEA?

Sakana AI has stated these models are not available in the EU/EEA. Availability elsewhere, including Japan, has not been explicitly addressed, so check official sources before use.

What is an orchestrator model?

Rather than a single large model answering directly, a trained orchestrator routes tasks to a pool of open-weight or specialized models, recursively calling itself as needed to assemble the final answer.

Are the benchmark numbers independently verified?

All figures in this article come from Sakana AI's own announcements. As of September 2026, no independent third-party verification has been found.

Feel free to contact us

Contact Us