TLDRocket
Sign in

Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text

MarkTechPost Asif Razzaq

Cloudflare launched Clef and Clef-flash, two open-weight decision models that answer typed questions with probabilities. They already run on Workers AI and can be self-hosted.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Cloudflare has shipped Clef and Clef-flash, its first models from the Workers AI team. They are not chatbots. Instead of generating a free-form reply, they take an input state plus a schema of typed questions and return probabilities for the allowed answers. That makes them feel less like a text generator and more like a machine for structured decisions.

The pair is open-weight under Apache 2.0, and both are compatible with TypeSafe AI’s Jev API. Switching from Jev is described as a simple endpoint-and-model-name change. Cloudflare says both models are live on Workers AI now, and the weights are also on Hugging Face for self-hosting.

Clef supports three question types: noul for yes-or-no answers, choice for selecting one named option with per-option probabilities and a confidence value, and score for rating against an ordered rubric. On Workers AI, a single request can include up to 64 questions and up to 4 images. The models also use a 64,536-token context window, which is plenty of room for messy real-world inputs.

Under the hood, Clef is post-trained from Qwen3.8-27B and Clef-flash from Qwen3.5-9B. Cloudflare froze both backbones and trained a schema-routing head on top, using rank-256 low-rank adapters plus a mix of label-smoothed cross-entropy and Brier loss. It also added a secondary training objective called Reinforcement Learning for Calibrated Decisions, which gives partial credit for adjacent ordinal choices.

Cloudflare’s own benchmark table shows the split personality clearly. On a 10-benchmark shortlist from the Decision Index 0.2.1 suite, a Clef model led on 7. It did especially well on BANKING77 and CLINC150+OOS, while Jev kept the upper hand on knowledge-heavy tests like GPQA Diamond, MMLU-Pro and BBH. In TypeSafe’s own workflow evaluations, Clef edged Jev in 3 of 4 areas, but only by small margins. None of the numbers have independent replication yet.

My take — AI-written commentary, not fact-checked reporting

This is the rare AI release that sounds like it was built for paperwork instead of demos, which is refreshing. The industry has spent long enough pretending every problem wants a chatbot; sometimes you just want a calibrated answer that doesn’t improvise. Cloudflare going after structured decisions feels more honest than another parade of eloquent nonsense.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.