
On October 1, Cloudflare released its first in-house trained AI models, and they cannot do something unusual: write text. Clef and Clef-flash only answer predefined questions and return a probability for every allowed answer. That sounds modest, but it targets exactly the point where many AI agents lose time and money today. Because the weights are available under a permissive license, anyone can download the models, run them and adapt them.
Key takeaways
- Clef has 27 billion parameters and Clef-flash has 9 billion; both are built on Alibaba’s Qwen models and are available on Hugging Face under the Apache 2.0 license.
- Instead of text, they return probabilities for yes/no questions, multiple-choice questions and rating scales, with up to 64 questions and 4 images per request.
- According to Cloudflare, median latency is 209 milliseconds for Clef and 39 milliseconds for Clef-flash; Clef leads on 7 of 10 classification benchmarks.
- On knowledge and reasoning tasks, Jev, the model from startup TypeSafe that inspired Clef, is clearly stronger, and Jev is cheaper per token.
- For developers, Clef pays off mainly where agents constantly make small decisions: sorting requests, checking content, detecting bots.
What makes a decision model different
A large language model such as Claude or GPT generates its answer word by word. If a program only needs to know whether a support ticket is urgent or which team should handle it, that is cumbersome: the model writes a sentence, the program has to parse it, and the same question may get a different answer the second time. Decision models flip this around. The developer passes in a state, such as the text of a ticket, a JSON object or an image, along with a schema of typed questions. What comes back is a number between zero and one for every valid answer.
Clef supports three question types: yes/no questions with the probability of yes, multiple-choice questions with a probability for each option, and ratings on an ordered scale. Technically, the underlying Qwen model reads the state and the questions in a single pass without generating a single token. A small additional transformer, which Cloudflare calls the joint schema head, computes the answer probabilities from that. This explains the speed: the slowest part of a language model, writing step by step, is gone.
The idea did not originate at Cloudflare. In mid-September, San Francisco startup TypeSafe AI introduced Jev, a model of this kind, and dubbed the category System One models, a nod to Daniel Kahneman’s fast, intuitive mode of thinking. Cloudflare also borrows the name of the training method meant to make the probabilities reliable: Reinforcement Learning for Calibrated Decisions. Clef is explicitly compatible with Jev’s API, so anyone already using Jev can compare the two without rewriting code. OpenAI also announced a Decisions API based on GPT-6 Luna in limited preview at its DevDay, but has not yet published pricing or details.
The numbers: fast at sorting, weak on knowledge
Cloudflare compares Clef with Jev and other models on ten classification benchmarks. Clef leads on seven of them. On BANKING77, which assigns banking queries to 77 categories, it reaches a macro-F1 score of 94.20, compared with 79.74 for Jev. The score measures how accurately a model classifies across all categories. On CLINC150, which also has to recognize when a query fits no category, the result is 97.43 to 89.27. Cloudflare puts median latency at 209.3 milliseconds for Clef, 38.8 milliseconds for Clef-flash and 524.1 milliseconds for Jev.
Where Clef loses is just as revealing. On tasks that require expert knowledge and multistep reasoning, Jev is far ahead: 82.7 to 65.9 on the MMLU-Pro knowledge test and 78.3 to 48.0 on the GPQA Diamond science test. Clef is a sorter, not a thinker. For decisions that require expert knowledge, a large language model is still the better choice.
Vendor numbers are one thing. Developer Flavio Copes tried the models himself on launch day and, testing from Italy, measured median latencies of 191 to 205 milliseconds for Clef-flash and 524 to 726 milliseconds for Clef, including the network round trip. Requests with images took between 13 and 30 seconds, and he had to shrink large images first. That puts the lab figures in perspective, but it does little to change the basic finding: for text classification, Clef-flash is fast enough to be built into almost any workflow.
What Clef costs and who it is for
On Cloudflare’s Workers AI platform, Clef costs $0.24 per million input tokens and Clef-flash $0.09. Jev is much cheaper at $0.042. Copes worked through an example: 100,000 ticket classifications cost about $1.68 with Jev and about $9.60 with Clef. Either way, those are tiny sums compared with the cost of the same work on a large language model. For agents that check many intermediate steps, that adds up quickly, as the heavy token consumption of today’s top models shows, something we described in our look at Gemini 4 Argon.
The real difference is the license. Jev and OpenAI’s Decisions API run only as hosted services. Clef, by contrast, is released under Apache 2.0, a license that also permits commercial use and modification. Organizations that cannot send sensitive data such as customer requests or internal documents to a cloud service can run the models in their own data center. Cloudflare tested them on a single Nvidia H200 GPU. By our calculation, Clef’s weights alone take up about 54 gigabytes at 16-bit precision, and Clef-flash’s about 18 gigabytes. That makes the smaller version interesting for well-equipped workstations, but it is not something for a laptop.
Cloudflare cites its own internal uses: the threat intelligence team classifies websites and spots phishing along the way, support routes requests to the right team by urgency, and Clef helps tell harmless crawlers from malicious bots. In a domain classification workflow, Clef took 2.2 seconds according to Cloudflare, compared with 4.7 seconds for the open language model GPT-OSS-120B. Because Clef can also read images, screenshots can, for example, be checked before publication for readable passwords or access keys. Nvidia’s approach with OpenShell and Sentry also shows how important automated checkpoints like these are becoming for agents.
How to try Clef
The fastest route is Workers AI, where the models are called @cf/cloudflare/clef and @cf/cloudflare/clef-flash and can be invoked from a Cloudflare Worker or through the REST API. If you want to self-host, the weights are on Hugging Face under cloudflare/clef and cloudflare/clef-flash and run with PyTorch and the Transformers library. A good first experiment is a task that a language model handles today with a long prompt, such as sorting incoming email: define three questions as a schema, run a few hundred real examples and compare the probabilities with your current results. If the confidence of an answer falls below a threshold you set, the program can hand the case off to a larger model or a human. That is exactly where calibrated probabilities are more useful than a free-form sentence.
For companies with their own data, Cloudflare is also announcing a service for tailoring Clef to a specific task with reinforcement learning, meaning training through rewards. Cloudflare is starting with selected design partners, and a self-serve platform is supposed to follow later. Cloudflare itself notes that such fine-tuning can weaken general capabilities in favor of the specialized task.
Outlook: small models take over the routine work
Clef is not a competitor to Claude, GPT or Gemini but a building block underneath them. Much of what agents do consists of routine decisions: Is this email relevant, does this request need a human, can this step run without asking first? Until now, those questions have often been answered by the same expensive language model that does the actual work. The fact that a startup, OpenAI and Cloudflare all arrived at the same idea within a little over two weeks suggests a distinct class of models is emerging. Cloudflare’s contribution is openness: anyone who wants their agents to make decisions that are traceable, fast and run on their own hardware now has a freely usable tool for it for the first time. Whether the probabilities hold up as well in practice as Cloudflare promises will be decided by real data, not benchmarks.
Sources
- Cloudflare Blog: Clef decision models
- Cloudflare Changelog: Introducing Clef on Workers AI
- Hugging Face: cloudflare/clef-flash
- MarkTechPost: Cloudflare releases Clef and Clef-flash
- Flavio Copes: A deep dive into Clef, Cloudflare's decision model
- Startup Fortune: Cloudflare open sources Clef
- Runtime Wire: TypeSafe opens Jev early access

