On September 11, 2026, Sakana AI launched Fugu Max v1.0 and Fugu Ultra v2.0, effectively transforming multi-agent orchestration into a standardized, API-compatible product. By pricing Fugu Max at $2 per million input tokens and $6 per million output tokens, Sakana has undercut the output costs of frontier models like Sonnet 5, GPT 5.6 Terra, and Kimi K3 by 40-60%. This move signals a shift in the AI pricing wars, moving the battlefield from individual model inference costs to the economics of orchestration.
What that actually means for builders is a fundamental decoupling of performance from proprietary lock-in. Fugu Max and Fugu Ultra are not monolithic models; they are orchestration engines. Their architecture is built on learned model coordination, detailed in two ICLR 2026 papers: TRINITY, which utilizes an evolved LLM coordinator, and Conductor, which employs reinforcement learning to discover natural-language coordination strategies. By routing tasks through a swappable pool of open-weight and specialized models, Sakana avoids the dependency on any single proprietary frontier model.
The technical performance of these systems is significant. Fugu Max achieves the best overall score on six benchmarks, including Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and SWEFish. Meanwhile, Fugu Ultra achieves best or joint-best results on five of eight benchmarks, including GDP.pdf, Chartography, DeepSWE, Toolathon, and SWEFish. Fugu Ultra is priced at $5 per million input tokens, $30 per million output tokens, and $0.50 per million cached input tokens, with a premium tier of $10/$45/$1.00 for contexts exceeding 272K tokens.
The part that gets hidden in the excitement over benchmark scores is the structural shift in pricing power. When the orchestration layer becomes model-agnostic, the value migrates away from the model providers and toward the orchestrators. If an orchestrator can deliver frontier-level performance by dynamically routing to cheaper, specialized models, the underlying model providers lose their ability to command premium margins. This mirrors the cost-cutting trends seen in DeepSeek V4.1 Flash, where architectural efficiency dictates market viability.
This development challenges the compute landlord thesis-the idea that those who own the infrastructure capture the most value. Instead, value is increasingly captured by those who control the traffic. If Sakana can maintain this 40-60% price advantage while delivering superior reasoning, the incentive for enterprises to commit to a single-model API provider evaporates. By offering a product that integrates with partners like OpenRouter, Vercel, opencode, Creao, and Merge, Sakana is positioning itself as the intelligent middleware of the agentic era.
The implications for the ecosystem are stark. Sakana has also introduced Fugu Cyber, a specialized variant for cybersecurity reasoning tasks, which achieves 86.9% on CyberGym and 72.1% on CTI-REALM. This demonstrates how orchestration can be tuned for specific domains without requiring a massive, general-purpose model for every query. They are not just selling tokens; they are selling the efficiency of the workflow.
What to watch next is whether the major labs respond by commoditizing their own orchestration layers or by attempting to tighten their API ecosystems. If the orchestration layer becomes the primary interface for developers, the model providers risk becoming mere utility suppliers, stripped of their brand power and pricing leverage. The era of the model-as-a-product is fading; the era of orchestration-as-a-product has arrived.
