,

Cloudflare releases its first open-weight models, and they do not write text

Cloudflare has released the first two models trained by its own Workers AI team. Clef has 27 billion parameters and Clef-flash has nine billion, both with a 64,000 token context. Neither writes prose. You hand the model a situation plus a schema of typed questions, up to 64 of them per request, and it returns a probability for each allowed answer: yes or no, pick one option, or score against a rubric. Clef can also take up to four images.

That is a narrower job than a chatbot does, and the point is what the narrowness buys. Routing a ticket, scoring urgency or deciding whether a message breaks a rule are all decisions with a fixed set of answers, and getting them back as probabilities rather than as sentences removes the parsing step and the room for the model to wander. Cloudflare quotes median latencies of 209.3 ms for Clef and 38.8 ms for Clef-flash.

The weights are on Hugging Face under Apache 2.0, so the models can be run outside Cloudflare, and the API follows an existing specification so current integrations can switch by changing the endpoint. Two caveats. The latency and benchmark figures are from Cloudflare’s own runs, not an independent evaluation. And the announcement gives no pricing, pointing only at the general Workers AI pricing page.

Source: Cloudflare, Introducing Clef: Cloudflare’s first open-source decision models, now on Workers AI.


Related