On September 1, the Silicon Data LLM Token Expenditure Index hit $0.97 per million tokens—the first time the industry benchmark has dipped below the $1 threshold since its inception. The index sat at less than half its summer peak, down 8.6% over the prior seven days. That number is the structural signal. What caused it was the most concentrated pricing compression event of the year: three frontier labs cutting costs within a 72-hour window.
On September 1, Anthropic launched Fable 5.1, headlined by a 75% reduction in cache-read pricing—from $1.00 to $0.25 per million tokens. The company estimates roughly 25% savings for typical workloads and up to 45% for the kind of context-heavy, tool-intensive agentic tasks that define the emerging agent economy. Base input and output rates remain at $10 and $50—the cut targets the repeat-read pattern that dominates long-running agent workflows.
The following day, Google introduced Gemini 3.8 Flash at introductory pricing of $0.75 and $3.75 per million tokens, valid through the end of 2026. The standard rate doubles to $1.50 and $7.50 on January 1—a temporary discount designed to capture volume during the current adoption window. Meta followed the same day with Muse Spark 1.3, which continues to offer a contributor tier at approximately $0.10 per million tokens for data-sharing partners, while holding its standard tier at $1.25 and $4.25.
These moves did not arrive in isolation. The compression wave began in late July, when OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%. Anthropic followed on August 10 by making the $2 and $10 pricing for Sonnet 5 permanent, canceling a scheduled September increase to $3 and $15. The $2 input tier has become the active war zone: Sonnet 5, GPT-5.6 Terra, and Gemini 3.1 Pro all sit at exactly that price point.
But the race to the bottom is only half the story. A dual-track structure is emerging, and the second track is where the real margin protection lives. While everyday models commoditize, the most capable—and most dangerous—models are being pulled off the public pricing grid entirely. OpenAI’s GPT-6 Astra, launched September 3, commands $10 and $50 standard pricing—the same as Anthropic’s Fable class—but its most consequential capabilities sit behind a separate, non-public tier for trusted defenders. Anthropic has restricted Mythos 5.1 to its Cyber Verification Program and Life Sciences Verification Program, gating access to the model’s strongest cybersecurity and biology capabilities. Google followed the same pattern, gating Gemini 3.8 Flash Cyber behind its new Fairwind Program for trusted defenders—not publicly priced.
The pattern is deliberate: commoditize the everyday tier to capture volume, then wall off the most powerful capabilities behind access-controlled programs that command whatever the market will bear. Roughly 95% of enterprise AI usage still runs on frontier models, according to Silicon Data—that volume flows through the commodity tier. The remaining 5%—the work that requires cyber-grade capabilities—flows through the gated tier at premium rates.
Not every player is following the compression script. DeepSeek V4-Pro-0813 bucked the trend in mid-August, raising prices by up to 14x to $1.32 and $3.96 per million tokens—a counter-signal that suggests the market is not universally racing downward. DeepSeek’s move likely reflects confidence in its model’s capability at the new price point, or a deliberate choice to prioritize margin over volume ahead of its own IPO preparations.
The primary driver for this bifurcation is financial. Both OpenAI and Anthropic filed confidential IPOs this summer. As these companies transition from research-heavy entities to public-market-ready businesses, they need two things simultaneously: volume metrics that justify scale, and margin stories that justify valuation. The dual-track structure delivers both—commodity pricing drives adoption numbers, while gated access protects the revenue per token that underwriters will scrutinize.
For builders, the takeaway is structural: the era of uniform frontier pricing is over. The market is splitting into a high-volume, low-margin utility layer and a high-value, restricted-access layer. The labs that can operate both tracks simultaneously—compressing commodity costs while gating their most powerful capabilities—will define the economics of the agentic economy. As our prior coverage of Anthropic’s triple release and the emerging revenue-share models from Chinese labs suggests, the ability to navigate this dual-track environment will determine which labs survive the transition from private research shops to public companies.
What to watch: whether the commodity tier stabilizes above the cost of goods or continues compressing toward zero, and whether the gated programs scale beyond government-aligned cybersecurity into broader enterprise adoption.
