Silicon Data’s Token Index Falls to 97 Cents, Tightening the Squeeze on Model Pricing
The benchmark’s new low points to cheaper model use for customers, while testing whether foundation-model providers can protect revenue as market rates fall.
Listen to this story
The audio brief
Story brief
3 key pointsSilicon Data’s token-cost benchmark dropped to 97 cents per million tokens Monday, below $1 for the first time and down 8.6% in a week. The usage-weighted index, based on traffic through multi-provider routing gateways, was below half its summer peak, reflecting cheaper Chinese models, dynamic pricing, and lower production costs. Users may get cheaper high-volume workloads, but providers face weaker pricing power...
- 01
The index is usage-weighted across selected models and routing gateways, so it is not a universal price for every provider or workload.
- 02
Moonshot’s Kimi K3 and other lower-cost Chinese models were cited as contributors to the decline.
- 03
OpenAI cut GPT-5.6 Luna pricing 80% and Terra pricing 20% on July 30.
AI providers have promised more capability for every dollar. A new market reading shows the commercial cost of delivering on that promise: Silicon Data’s LLM Token Expenditure Index fell to 97 cents on Monday, its lowest level since the measure began late last year. Lower rates can reduce customer bills while putting pressure on the providers that sell metered model access.
The index is a usage-weighted effective price for one million large-language-model tokens across a defined group of models. It combines provider prices with consumption volumes observed through multi-provider routing gateways, so it captures an aggregate market rate rather than a universal price for every model or workload.
Cheaper alternatives are part of the move
The decline has been attributed in part to lower-cost Chinese models, including Moonshot’s Kimi K3. Other cited pressures include dynamic pricing that changes access rates with demand and falling costs to produce tokens. Silicon Data research chief Steve Hou said the combined supply of frontier and cheaper models may already be capable enough for most tasks.
OpenAI’s July 30 price reductions are an earlier example of lower API rates at a major proprietary provider. The company cut GPT-5.6 Luna prices by 80% and GPT-5.6 Terra prices by 20%, saying it was passing efficiency gains through to customers. Starting that day, Luna was listed at $0.20 per million input tokens and $1.20 per million output tokens; Terra was listed at $2 and $12, respectively.
Lower bills change the provider equation
For model users, lower token rates can make more queries and high-volume workloads cheaper to run. For foundation-model providers, they can reduce pricing power and revenue. The exposure is especially acute where revenue falls while compute commitments remain fixed, according to Charles-Henry Monchau of Syz Group. The effect will still vary by provider, its prices and its usage mix; the index does not determine any company’s individual economics.
Sources
- cnbc.comAI token prices are hitting new record lows
- qz.comAI token prices hit record low, pressuring OpenAI and Anthropic
- openai.comAdvancing the price-performance frontier with GPT-5.6