Leiolai Launches Device-Run AI With 11M-Token Context and Output From $0.02 Per Million
The app and API depend on users supplying computation, making coordination, latency and trust central tests of its alternative to centralized AI infrastructure.
Listen to this story
The audio brief
Story brief
3 key pointsLeiolai is replacing its Early Access model with leiolai-1, a device-distributed AI service available through a free consumer app and OpenAI-compatible API. Its headline economics are $0.01 per million input tokens and $0.02 output for Fast and Research, while Private starts at $20/$90 and Plus Ultra reaches $160/$700. The model claims an 11-million-token context and uncapped generation, but the network’s promise...
- 01
An 11-million-token context window is offered in both app and API; Leiolai calls it the largest publicly available.
- 02
Users can earn by contributing phones, tablets, or computers, with estimated water savings shown after responses.
- 03
Private costs 2,000× Fast input pricing and 4,500× output pricing, while using trusted hardware and users’ own devices.
Leiolai has launched leiolai-1 with a consumer app and developer API built around an unusual infrastructure promise: model inference runs across users’ existing devices rather than centralized data centers. The company is pairing that design with payments for people who contribute computation, putting the size and dependability of its device network at the center of the product.
The launch arrives with a claimed 11-million-token context window, available in both the consumer product and API. Leiolai calls it the world’s largest publicly available context window. It also says leiolai-1 can continue generating without fixed output limits because runtime scales linearly as context grows.
That setup changes what expansion means. Instead of adding capacity through a new server build-out, Leiolai says its network gains computation as people add phones, tablets and computers. The app shows estimated water savings and user earnings after answers; the company also says its system uses no water for inference.
A familiar API over an unfamiliar compute layer
For developers, Leiolai is using an OpenAI-compatible Chat Completions API format. That compatibility is intended to reduce the work of moving software already built around that format, while leaving Leiolai’s distributed inference system underneath the request.
The company splits the service into modes that pair different workloads with different control and pricing choices. Fast and Research start at the same base rates, while Private is priced far higher. Developers can also set response depth, choosing how much computation to use for a request.
The modes divide bulk generation from confidential work
- Fast is aimed at high-output tasks such as synthetic datasets, evaluation corpora, document transformation, catalogue generation and output-heavy product features.
- Research is intended for public problems including protein design, mathematics, algorithm discovery and AI-related work.
- Private is designed to process requests confidentially and exclusively on trusted hardware and a user’s own devices. It starts at $20 per million input tokens and $90 per million output tokens.
- Plus Ultra, a higher-depth option in Research or Private, costs $160 per million input tokens and $700 per million output tokens.
The consumer app is free and does not require a payment card, Leiolai says. Users can earn money by contributing device computation, and the company says access to higher intelligence levels increases as users progress. The release replaces Leiolai’s earlier Early Access model with leiolai-1.
Leiolai claims a network of 10 million devices could provide more long-context throughput than all major frontier labs combined. That is a company projection, not a disclosed operating result. Reliability, coordination, latency and trust remain technical and operational challenges for distributed inference across consumer hardware.
Sources
- itbrief.newsLeiolai launches AI model that runs on user devices