Liner Details Model API Prices and Claims More Than 50% Internal Token Savings

The company’s August spending fell against its first-half baseline, it says. Customers’ savings will depend on the requests its router handles.

By 2 min read
Liner Details Model API Prices and Claims More Than 50% Internal Token Savings
Liner Details Model API Prices and Claims More Than 50% Internal Token Savings

Listen to this story

The audio brief

About 1:28
0:001:28
Read transcript
Liner says its August token expenses were more than 50 percent below its first-half 2026 baseline after it deployed a model-routing service. The company has now published prices for that service, called Liner Model API: one dollar per million input tokens, six dollars per million output tokens, and ten cents per million cached input tokens. The idea is to avoid sending every request to the same model. Liner’s Orchestrator estimates what each request needs, then routes it to one model—not several running in parallel. The company says that avoids the extra token costs of parallel calls. In an evaluation of requests to its own services, 43 percent went to higher-performance models and 57 percent to more efficient ones, with no loss in answer accuracy, Liner says. That split reflects Liner’s traffic, not a promise about what customers will see. There are two separate savings claims here. The August figure compares Liner’s expenses over time. Separately, the company estimates its API can cost at least 50 percent less than comparable models, based on published prices. Neither figure guarantees a customer’s savings; results depend on the requests, model choices, and quality requirements. Liner’s comparisons include Claude Sonnet 5 and GPT-5.6-Terra, but developers can test the routing against their own workloads. The deciding constraint is whether the selected models deliver acceptable answers for the work customers actually send.

Story brief

3 key points

Liner has put public rates on its Model API and reports that its Orchestrator cut its own August token expenses by more than 50% versus a first-half 2026 baseline. The API routes each request to one model, with Liner reporting a 43% high-performance/57% efficient-model mix on its own traffic without lower accuracy. That result is not a customer guarantee: its separate claim of at least 50% savings versus comparable...

  1. 01

    Published rates are $1 per million input tokens, $6 per million output tokens, and $0.10 per million cached input tokens.

  2. 02

    Liner’s 43%/57% routing mix came from requests to its own services; it is not a forecast for customer traffic.

  3. 03

    The API selects one model per request rather than querying several simultaneously, avoiding the extra token costs of parallel calls.

Liner has published prices and an internal savings result for its model-routing API, which it announced on September 15. The figures sharpen its case for replacing one default AI model with a service that selects a model for each request.

One request, one model

A fixed-model setup sends every request to the same model, including work a cheaper one might handle. Liner Model API instead weighs the expected quality and cost of candidate models, then selects one for each request. It does not call several models at once, a choice Liner says avoids the extra token costs of doing so.

Liner says harder coding, math and reasoning questions go to higher-end models, while simpler ones go to lighter models. In an evaluation of requests from its own services, about 43% went to high-performance models and 57% to more efficient ones without a loss in answer accuracy, the company says. That mix is not a promised split for other customers.

Published API rates
$1Input

Liner lists input at $1 per million tokens.

$6Output

Liner lists output at $6 per million tokens.

$0.10Cached input

Liner lists cached input at $0.10 per million tokens.

Two ways to read the savings claim

Liner reports that after deploying Liner Orchestrator, its August token expenses were more than 50% below its first-half 2026 baseline. Separately, it estimates the API can cost at least 50% less than comparable models in the same performance tier, based on publicly available prices.

Those are different comparisons: one tracks Liner’s expenses across time; the other compares its API with models it considers peers. Neither guarantees a customer’s savings. Liner says actual savings depend on the mix of requests, model use and workload requirements. Customers with many complex questions may see a different result from those with mostly routine ones.

The quality test sits inside the route

A lower bill is useful only if the selected model can do the work. Liner says it refined the router through real-world use and controlled benchmark tests. Its published comparisons include Claude Sonnet 5 and GPT-5.6-Terra. Developers still have to judge whether its model choices meet the quality their own applications require.

The company says it developed its evaluation system while running its own AI services and validated the product using de-identified, representative requests. Developers can create an API key through Liner’s platform, inspect its published benchmarks and use a calculator to estimate savings from their current model spending. The deciding test is whether lower costs come with acceptable answers on their own workloads.

Sources

  1. liner.com라이너 모델 API 출시 | LLM 비용 50% 절감 솔루션
  2. manilatimes.netLiner Launches Liner Model API to Cut Enterprise LLM Costs by More Than 50%

Loading discussion...

YOUR READING SPACE

Notifications