Cloudflare Releases Open Models for AI Decisions Without Text Generation
Clef and Clef-flash return probabilities that software can act on directly. Cloudflare’s tests show faster responses than Jev, but the smaller model’s accuracy varies sharply by task.
Cloudflare’s Clef and Clef-flash package classification and scoring into typed probability outputs that application code can act on directly, rather than text an agent must parse. Released October 1, 2026, they run on Workers AI and are downloadable under Apache 2.0, with a Jev-compatible interface intended to ease migration. Cloudflare reports substantially lower median latency than Jev, but results vary by task, so teams should validate accuracy before letting model scores trigger actions; customization initially requires Cloudflare engineers.
01
Requests can contain up to 64 yes/no, choice, or scoring questions, up to four images, and a 64,000-token context window.
02
Across Cloudflare’s 43 benchmark runs, median latency was 38.8 ms for Clef-flash, 209.3 ms for Clef, and 524.1 ms for Jev; the comparisons are company-run.
03
On CLINC150+OOS, Clef scored 97.43, but Clef-flash scored 66.77 versus Jev’s 89.27, showing the variants are not interchangeable.
An AI agent can route a support ticket or trigger an escalation without first waiting for a model to write an answer. Cloudflarereleased Clef and Clef-flash on October 1, 2026, to supply those decisions directly: predefined answers with probabilities, rather than free-form text.
Both models are available on Workers AI, Cloudflare’s hosted inference platform, and their weights are downloadable from Hugging Face under Apache 2.0. They also follow TypeSafe AI’s Jev interface: an existing integration can switch by changing its endpoint and model, according to Cloudflare.
The questions arrive with the request
A Clef request supplies the current state—such as a customer’s message—and typed questions that define the permitted answers. The model returns probabilities that downstream code can use immediately. Cloudflare’s product changelog says a request can contain up to 64 questions, across three formats:
Yes or no: a binary question returns the probability that the answer is yes, such as whether a support request is urgent.
A choice: developers define the options, and Clef returns the selected option, probabilities for each option, and a confidence value.
A score: an ordered rubric returns a probability-weighted score and probabilities for each level, allowing software to assess a submission against supplied criteria.
There is no free-form response to parse and no reasoning text to wait for. Underneath, Cloudflare says it keeps the Qwen base models frozen while training added routing components and adapters. Its training combines synthetic datasets, objectives for answer accuracy and probability quality, and a method called Reinforcement Learning for Calibrated Decisions.
The input is not limited to text. Requests can include up to four images, and both models support a 64,000-token context window—the amount of input they can hold. Cloudflare contrasts that with Jev’s text-only inputs and 32,000-token window, making visual classification another point of competition beyond speed.
Fast responses, uneven task results
Cloudflare’s headline performance claim comes from 43 benchmark runs. It reports Clef as about 2.5 times faster than Jev at the median, and Clef-flash as about 13 times faster. These are company-run comparisons, not independently verified production results.
The slower end of the latency distribution also differs from the median. Cloudflare lists 95th-percentile response times of 122.4 milliseconds for Clef-flash, 238.6 milliseconds for Clef, and 536 milliseconds for Jev. The 38.8-millisecond figure is therefore not a promise for every request.
On quality, Cloudflare says one of its Clef models scores highest on seven of 10 decision benchmarks. Its BANKING77 banking-intent classification results favor the larger model: Clef scores 94.20, Clef-flash 90.93, and Jev 79.74 on the test’s macro-F1 measure. But the family-level win count does not mean the two Clef variants are interchangeable.
On CLINC150+OOS, Clef scores 97.43, while Clef-flash scores 66.77, below Jev’s 89.27. Jev also leads on When2Call. The Decoder reproduces Cloudflare’s accuracy figures: 80.97 for Jev, 72.37 for Clef, and 65.58 for Clef-flash. These task-level gaps complicate any claim that faster decisions are uniformly better decisions.
A website check includes more than inference
Cloudflare’s threat intelligence team has been testing Clef on website domains. In one example, the workflow fetched, rendered, and classified a site in 2.2 seconds, assigning probabilities to categories including fashion, ecommerce, and phishing.
The same workflow took 4.7 seconds with gpt-oss-120b, which returned only two classifications. Unlike the millisecond benchmark figures, these times include obtaining and rendering the website.
Cloudflare argues that a person need not be involved in every agent decision. Its described workflow still allows code to defer to a human when needed. That distinction leaves an important choice with the application: whether a probability should trigger an action, an escalation, or review. Clef supplies the classification; downstream software determines what follows.
Customization starts with Cloudflare engineers
The release also includes a reinforcement-learning fine-tuning service for adapting Clef to customers’ workloads. Initially, Cloudflare engineers will work directly with customers; a self-service platform is planned for later. The customization offering is tied to this model family, rather than a separate general-purpose model launch.
The proposed pipeline uses AI Gateway to log requests for a dataset, containers to evaluate them in a training sandbox, and a new trainer component to deploy the customized model on Workers AI. Cloudflare also plans internal uses for abuse-report review, support triage, and distinguishing useful bots from harmful ones—workloads where the resulting classifications can feed operational decisions.
Editorial illustration for Cloudflare Releases Open Models for AI Decisions Without Text Generation.Source: the-decoder.com.
Sources
blog.cloudflare.comIntroducing Clef: our open-source decision models, and new RL fine-tuning platform
Reader comments
Newest comments first. Replies stay oldest first.