Cloudflare has released two decision models for agents, Clef and the smaller Clef-flash. A decision model does not write a long answer. It assigns probabilities to options that were set in advance. The Decoder’s example is a customer-support message: Clef judges how urgent it is and which team should handle it. Later code can use that result to route a ticket, trigger an escalation, or hand the case to a person.
Cloudflare’s line is that a human does not necessarily need to stay in the loop for an agent’s decisions. The same passage says an agent can gather context, decide, act, or defer to a human when needed. The two sentences belong together. The first is the company’s claim. The second says handing the case to a person is still part of the design. The name comes from a clef in music: a clef says which line is which pitch, and a decision model says which set of options the next action falls into.
[1]
The company places Clef between a large language model and a traditional classifier. A language model can reason and call tools, but its output varies and it can be slow. A traditional classifier is fast, but each new category means training again. Cloudflare keeps the interface compatible with TypeSafe AI’s Jev so existing customers can switch. The article says Jev was introduced in mid-September by TypeSafe, founded by former OpenAI researcher Diogo Almeida. TypeSafe describes Jev as producing no invented answers. The Decoder’s limit on that claim is narrow: the model stays inside the preset options, and that does not mean it picks the right one.
Speed is Cloudflare’s main selling point, and the numbers are its own. The company says that across 43 benchmarks, Clef and Clef-flash are faster than the competing decision models it treats as relevant. Median latency is about 39 milliseconds for Clef-flash and about 209 milliseconds for Clef, against just over 524 milliseconds for Jev. Both models run on Cloudflare’s own infrastructure. In Cloudflare’s own figures, Clef has the highest decision quality, and Clef-flash comes near Jev’s accuracy at a much lower latency.
[1]The accuracy table does not have Clef first on every row. The figures below are the ones in the article, not a rewritten ranking. On API Bank, Clef is 91.93, Clef-flash is 93.11, and Jev is 88.19. On When2Call, Clef is 72.37, Clef-flash is 65.58, and Jev is 80.97. On PhishNChips, Clef is 79.60, Clef-flash is 75.05, Jev is 62.55, and DiffusionGemma is 85.35. Jev is higher on When2Call. DiffusionGemma is higher on PhishNChips.
In one example from the threat-intelligence team, a domain is given a 95 percent probability of being a fashion site, 85 percent of being an online store, and under 1 percent of being a phishing site. Fetching, rendering, and classifying took 2.2 seconds. The company’s fastest general-purpose language model took 4.7 seconds for the same process and returned only two categories. That is Cloudflare’s example, not an independent retest. Clef can take images. The article says Jev is still limited to text. The context window is 64,000 tokens, twice Jev’s.
[1]Cloudflare says Clef is based on Qwen3.8-27B and Clef-flash on Qwen3.5-9B. The base models stay unchanged. The company trains extra components on its own synthetic data, and those components turn the model’s internal computations into options and probabilities. It also uses its own variant of the calibrated-decision training method TypeSafe used for Jev. That is a variant of the method, not Jev under a new name.
For customers, Cloudflare is also offering a reinforcement-learning service so they can fit Clef to their own tasks. At first, engineers working with the customer handle the fine-tuning. A self-serve platform is planned for later. The data can come from request logs in AI Gateway, be evaluated in containers used as a sandbox, and then be deployed to Workers AI by a new trainer. Custom models use technology from Replicate, which Cloudflare acquired in late 2025. Both models are on Workers AI and on Hugging Face under the Apache-2.0 license. The company plans to use them internally to review abuse reports, sort support requests, and separate useful bots from harmful ones. Those are planned internal uses.
The article mentions in one sentence that OpenAI, in late September, offered a Decisions interface on GPT-6 Luna that accepts text or images. That is another company’s product and is not developed here. Clef’s latency, the 43 benchmarks, and the table scores are what Cloudflare reports. The Decoder does not provide an independent measurement.
[1]要点
- Clef and Clef-flash assign probabilities to preset options instead of writing long text. Cloudflare says a human need not stay in the loop, and the same passage says the agent can defer to a person.
- Self-reported median latency is about 39 and 209 milliseconds, against just over 524 for Jev. The bases are Qwen3.8-27B and Qwen3.5-9B.
- On API Bank, Clef-flash at 93.11 is above Jev. On When2Call, Jev at 80.97 is higher. On PhishNChips, DiffusionGemma at 85.35 is higher.
- The weights are on Hugging Face and Workers AI under Apache-2.0. The latency and the table are Cloudflare’s own numbers.