Cloudflare released two open-weight models on Thursday that do not write anything. Clef and Clef-flash take a question and a list of options, then return a probability for each option, so the developer’s own code decides what happens next. Cloudflare calls this category a decision model, and it is the company’s answer to Typesafe AI’s Jev, which has drawn heavy attention over the past few weeks.

The idea is easiest to see with a support ticket. Feed the model a message saying checkout has failed for every customer for an hour, and ask it two things: how serious the problem is, and who ought to own the fix. It hands back scores for each. Your software can then route the ticket, page someone, or send it to a human when the scores look unsure. A chatbot would produce a paragraph. This produces a number you can put in an if-statement.

Cloudflare, writing on its own blog, says the models are free to download on Hugging Face under an Apache 2.0 license, and are also hosted on its Workers AI platform with an interface that matches Jev’s, so a team can swap one for the other. Clef reads images as well as text, which the post says Jev does not, and its 64,000-token input window is double Jev’s 32,000.

The company’s own security team supplied the demo. Asked to categorize a website, Clef fetched the page, rendered it and returned a spread of scores (95 percent fashion site, 85 percent ecommerce, under 1 percent phishing) in 2.2 seconds. Cloudflare says its fastest general-purpose model, gpt-oss-120b, needed 4.7 seconds for the same job and returned only two labels.

Every number in the post comes from Cloudflare’s own testing, and the headline claim is that Clef leads on the Jev Decision Index, a benchmark suite named for the rival it is chasing. The company’s results table is more mixed than that framing suggests. Jev scores higher on the When2Call test of whether a model knows when to ask for help instead of acting, and also on the BRIGHT retrieval test. On a phishing test, a DiffusionGemma-based entry that Cloudflare also lists beats Clef. Clef’s wins include banking-intent and intent-detection tests, where the margin over Jev runs from roughly 8 to 15 points.

On Typesafe’s own four-workflow suite, Clef beats Jev in three: invoice processing, customer service and security incidents. The customer service margin is 0.3 points. Jev wins the fourth, tracing what an agent did, by about three points. Cloudflare reports that the smaller Clef-flash outscores the larger Clef on several tests and fares well on the workflows given how fast it runs.

Speed is the real pitch. Because the model scores options in one pass instead of generating text word by word, Cloudflare reports a median response of about 209 milliseconds for Clef and 39 milliseconds for Clef-flash, against 524 milliseconds for Jev. The post concedes that a model called Laya is faster still but weaker on quality. Hosting on Cloudflare’s edge network trims network delay on top, which is why the company argues a decision model can sit inside an agent’s hot path.

Underneath, Cloudflare built Clef on frozen Qwen models (a 27-billion-parameter version for Clef, a 9-billion version for Clef-flash) with small trained adapters and a custom scoring component on top. It trained for calibration, meaning the stated probabilities should match how often the model is actually right, using a Brier loss, a standard penalty for overconfident guesses. It added a reinforcement-learning stage the company calls RLCD, which gives partial credit for near-miss answers.

Cloudflare also announced a fine-tuning service. At first it is hands-on, run by the company’s forward-deployed engineers, with a self-serve platform planned later. That platform would stitch together AI Gateway, which logs a customer’s traffic into a training set, plus Containers and a new trainer component.

The commercial logic is clear. Decision models are small, cheap and open, so margins are thin, and Cloudflare is selling the surrounding plumbing: hosting, request logs, retraining. Teams already testing Jev should run Clef on their own labeled data before trusting anyone’s index, since a gap of a point or two on a vendor’s chosen benchmarks will not survive contact with a real ticket queue.

Reported by Cloudflare’s blog on 1 October 2026.