Cerebras’s fast dinner-booking agent had a head start: it already knew how to navigate the website. In a September 24 demonstration, the company reported a median completion time of 22 seconds across two successful attempts. Faster computing helped, but so did instructions prepared before timing began.
Cerebras built the assistant with Qwen 3.8 27B running on its computing infrastructure and Pi as the agent harness. A harness is the software that manages the conversation, runs tools, handles errors and sends results back to the model. Cerebras added booking tools to Pi.
A browser agent receives a page result, asks the model what to do next and executes another action. Each return to the model adds waiting. An assistant can also spend time discovering a website’s controls or checking options one after another.
Three changes to the booking route
- Check restaurants together. Cerebras directed independent restaurant checks to run in parallel. The booking decision still depended on their results, but gathering those results no longer required finishing one check before starting another.
- Shorten model waits. Faster inference—the process of generating a model’s response—reduced pauses between actions. It cannot make a slow website load instantly or eliminate phone verification. Some actions still require an earlier result.
- Reuse the navigation procedure. The team converted its knowledge of the booking site into instructions the agent could reuse. Cerebras said this reduced tool calls by more than 80%, avoiding repeated discovery during the timed task.
The saved skill did not replace changing information with stored answers. Restaurant availability, prices and whether a card was required still needed live checks. Preparation covered how to find that information, not permission to assume it remained correct.
The work of figuring out the route happens before the timed run; it doesn’t disappear.
Sarah Chieng and Sherif Cherfa, Cerebras
Cerebras sent the same reservation request to Meta Muse, Claude Cowork and Grok Bot. Their successful recorded runs took 4 minutes 36 seconds, 6 minutes 25 seconds and 7 minutes 40 seconds, respectively. The company also cited a 37-second manual booking.
The assistants took different routes. Meta Muse made nine direct OpenTable API calls; Claude Cowork made 57 tool calls. One Grok browser run spent 2 minutes 18 seconds checking three restaurants sequentially.
Cerebras calls its result 19 times faster than existing assistants, while labeling the examples “Recorded runs, not a general ranking.” Its explanation says the optimized comparison bundles changes to the model, harness and execution path; it does not isolate the saved skill’s contribution. How well the assistant handles other errands remains unresolved.
Reader comments
Newest comments first. Replies stay oldest first.