Gimlet Adds Cerebras Chips to Its AI Cloud, With a Data Center Planned
An integrated system is already serving private deployments. The companies still have to deliver the planned data center and substantiate their speed target.
Loading page…
An integrated system is already serving private deployments. The companies still have to deliver the planned data center and substantiate their speed target.
Listen to this story
Gimlet and Cerebras are moving beyond separate hardware options toward a cloud service that routes phases of a single AI inference request across Cerebras chips and GPUs. The system is already serving tokens in private deployments, but the companies have not published performance results; Gimlet’s goal of up to 3,000 tokens per second remains a target without a specified model or test setup. A Cerebras-powered Gimlet Cloud data center is planned for later in 2026, the step that could bring the service beyond private traffic.
The companies say joint customer work began in 2025, before their September 28 collaboration announcement.
Cerebras was named a launch partner for its CS-4 system; the partners also plan joint work on infrastructure, software, developer tools, testing, and operations.
The partners point to voice, video, assistants, and agents as use cases where lower response latency may matter, but have not demonstrated performance across them.
Gimlet Cloud has begun serving AI requests using a combination of Cerebras hardware and GPUs, though only in private deployments so far. Gimlet Labs and Cerebras announced the collaboration on September 28, with a Cerebras-powered data center expected later in 2026. Their bet is that assigning different parts of an AI response to different chips can make the service faster at production scale.
Inference is the work of running an AI model to produce a response. The companies’ design uses Gimlet’s software to direct phases of that work to the hardware they say suits each phase best. Cerebras supplies its Wafer Scale Engine, while GPUs remain part of the same serving system. This is a plan for combining chip types within a workload, not simply offering developers two separate hardware choices.
The division reflects how the partners describe their respective strengths. Cerebras co-founder and CTO Sean Lie says its systems produce fast tokens—the pieces of text an AI generates—while GPUs provide high throughput. Gimlet’s job is to coordinate those resources through its cloud rather than require developers to manage the underlying chips. Whether that division delivers both quick responses and enough overall output is the central test of the design.
The collaboration did not start with this announcement. Gimlet and Cerebras say they have worked together on customer engagements since 2025, and that an integrated system is already serving tokens in private deployments. That establishes some use beyond a proposed architecture, but the announcement does not give a public performance result for those deployments.
The planned data center is a different step from serving private traffic today. To bring the combined service more broadly into Gimlet Cloud, the companies say they will extend their work across infrastructure design, software, developer tools, testing and production operations. Cerebras also named Gimlet a launch partner for its CS-4 system, giving the cloud provider a planned route to that hardware.
Gimlet CEO Zain Asgar says the combined system is intended to deliver up to 3,000 tokens per second at production scale. The 3,000-token figure is a target, not a published test result. The announcement does not specify a model or measurement setup for it, leaving readers unable to judge how that speed would translate to a particular application.
The intended audience helps explain the emphasis on speed. The partners point to voice, video, assistants and AI agents as applications where waiting for a response can disrupt an interaction. Faster token generation could help those uses, but the companies have not shown in this announcement how the integrated system performs across them. A production deployment would put that promise in a setting beyond the private deployments described so far.
Loading discussion...
Join the conversation
What would make a speed promise convincing to you?
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.