Businesspublished

Silicon Data’s Token Index Falls to 97 Cents, Tightening the Squeeze on Model Pricing

The benchmark’s new low points to cheaper model use for customers, while testing whether foundation-model providers can protect revenue as market rates fall.

By 2 min read
Silicon Data’s Token Index Falls to 97 Cents, Tightening the Squeeze on Model Pricing
Silicon Data’s Token Index Falls to 97 Cents, Tightening the Squeeze on Model Pricing

Listen to this story

The audio brief

About 1:28
0:001:28
Read transcript
Silicon Data’s token-cost index has fallen below one dollar for the first time, reaching 97 cents per million tokens on Monday. That is an 8.6 percent drop in a week, and less than half the benchmark’s peak earlier this summer. For customers running large volumes of model requests, that points to cheaper usage. For the companies selling access, it raises a harder question: how much pricing power is left as model rates keep falling? The index is a usage-weighted effective price across a defined set of large language models. It combines provider prices with real consumption seen through multi-provider routing gateways, so it reflects an aggregate market rate—not a universal price for every model or workload. Lower-cost Chinese models, including Moonshot’s Kimi K3, are part of the decline. Silicon Data research chief Steve Hou also pointed to dynamic pricing and falling production costs, saying the combined supply of frontier and cheaper models may already handle most tasks. OpenAI illustrated the trend on July 30, cutting GPT-5.6 Luna’s prices by 80 percent and Terra’s by 20 percent. Luna then cost 20 cents per million input tokens and $1.20 per million output tokens. Lower bills could expand high-volume use, but provider impact will vary with pricing, usage mix, and fixed compute commitments. The key constraint is whether revenue can fall faster than those commitments.

Story brief

3 key points

Silicon Data’s token-cost benchmark dropped to 97 cents per million tokens Monday, below $1 for the first time and down 8.6% in a week. The usage-weighted index, based on traffic through multi-provider routing gateways, was below half its summer peak, reflecting cheaper Chinese models, dynamic pricing, and lower production costs. Users may get cheaper high-volume workloads, but providers face weaker pricing power...

  1. 01

    The index is usage-weighted across selected models and routing gateways, so it is not a universal price for every provider or workload.

  2. 02

    Moonshot’s Kimi K3 and other lower-cost Chinese models were cited as contributors to the decline.

  3. 03

    OpenAI cut GPT-5.6 Luna pricing 80% and Terra pricing 20% on July 30.

AI providers have promised more capability for every dollar. A new market reading shows the commercial cost of delivering on that promise: Silicon Data’s LLM Token Expenditure Index fell to 97 cents on Monday, its lowest level since the measure began late last year. Lower rates can reduce customer bills while putting pressure on the providers that sell metered model access.

The index is a usage-weighted effective price for one million large-language-model tokens across a defined group of models. It combines provider prices with consumption volumes observed through multi-provider routing gateways, so it captures an aggregate market rate rather than a universal price for every model or workload.

Cheaper alternatives are part of the move

The decline has been attributed in part to lower-cost Chinese models, including Moonshot’s Kimi K3. Other cited pressures include dynamic pricing that changes access rates with demand and falling costs to produce tokens. Silicon Data research chief Steve Hou said the combined supply of frontier and cheaper models may already be capable enough for most tasks.

OpenAI’s July 30 price reductions are an earlier example of lower API rates at a major proprietary provider. The company cut GPT-5.6 Luna prices by 80% and GPT-5.6 Terra prices by 20%, saying it was passing efficiency gains through to customers. Starting that day, Luna was listed at $0.20 per million input tokens and $1.20 per million output tokens; Terra was listed at $2 and $12, respectively.

Lower bills change the provider equation

For model users, lower token rates can make more queries and high-volume workloads cheaper to run. For foundation-model providers, they can reduce pricing power and revenue. The exposure is especially acute where revenue falls while compute commitments remain fixed, according to Charles-Henry Monchau of Syz Group. The effect will still vary by provider, its prices and its usage mix; the index does not determine any company’s individual economics.

Sources

  1. cnbc.comAI token prices are hitting new record lows
  2. qz.comAI token prices hit record low, pressuring OpenAI and Anthropic
  3. openai.comAdvancing the price-performance frontier with GPT-5.6