Sakana AI has released Fugu Max and Fugu Ultra v2, two new models in its Sakana Fugu family. Fugu is not a single foundation model; it is a learned orchestrator that routes work across a pool of other models behind one API. The new release tunes that architecture for two missions: Fugu Max targets the best output per dollar, while Fugu Ultra v2 targets the highest capability on hard, multi-step tasks.
Is it deployable? Yes, as a hosted API. Both models are live today through Sakana's OpenAI-compatible API. There are no open weights to self-host, and Sakana does not offer the service in the EU/EEA.
Why Sakana Frames This as a Two-Axis Problem
Sakana's argument is direct: real workloads are judged on capability and cost together. Sending a simple data lookup to a multi-trillion-parameter model wastes money. A better system picks the cheapest machinery that can still solve the task.
The Sakana team describes this using the Pareto frontier. On that frontier, gaining quality costs more, and cutting cost loses quality. Fugu Max and Fugu Ultra v2 share one core orchestration architecture. Only the optimization target differs.
The release follows a fast cadence. Fugu entered beta in April, reached general availability in June, and added Fugu-Cyber and a Claude Code interface in July.
How Fugu Orchestration Works
The Sakana Fugu Technical Report describes [the orchestration mechanism in detail]. [Content truncatedβplease provide the remainder of the source to complete the article.]
via MarkTechPost
