AI

Cloudflare Clef decision models: $0.24/M vs Jev, open weights

· Geeknewz Author

Abstract glowing neural network visualization in blue and purple

Cloudflare trained its first Workers AI models and pointed them at a narrow job: make bounded decisions fast, without free-form chat. On October 1, 2026, the company launched Clef (27B) and Clef-flash (9B) on Workers AI, published Apache 2.0 weights on Hugging Face, and opened a design-partner lane for reinforcement-learning fine-tunes. The pitch is not another general LLM. It is a drop-in rival to Typesafe's Jev decision API, with vision and a longer context window attached.

That lands two weeks after Jev made structured yes/no, choice, and score questions feel like a real product category. Below is the price and latency math from Cloudflare's own docs, what open weights actually buy you on hardware, and who should swap an agent hot path this month.

Person typing on a laptop with code on the screen
Photo via Unsplash (https://unsplash.com/photos/1531482615713-2afd69097998). Free license.

What a decision model returns

Instead of generating paragraphs, Clef reads a state (text, JSON, images, or video) plus a schema of typed questions, then returns a probability for every allowed answer in one pass. Question types match the System One / Jev shape: noul for yes/no, choice for a labeled set, and score against an ordered rubric. You can ask up to 64 questions per request. Cloudflare's threat-intel example is concrete: feed a domain (with Browser Run), get category probabilities in about 2.2 seconds, versus 4.7 seconds and fewer labels from gpt-oss-120b in the same workflow.

ModelSizeWorkers AI priceContextBest for
Clef27B (Qwen3.8 backbone)$0.24 per M input tokens64KHighest-precision decisions
Clef-flash9B (Qwen3.5 backbone)$0.09 per M input tokens64KLatency-critical hot paths
Jev (Typesafe, per The Register)undisclosed$0.042 per M tokens32K state+question / 64K requestCheaper text-only decisions

Our math on hosted cost: $0.24 ÷ $0.042 ≈ 5.7× Jev's reported rate for full Clef, and $0.09 ÷ $0.042 ≈ 2.1× for Clef-flash. You are paying that premium for multimodal inputs (up to four images), a 64K context window on both Clef SKUs, and Cloudflare's claim that Clef leads seven of ten Decision Index-style benches it published. The Register notes those scores are still Cloudflare-reported, not yet reproduced as the official leaderboard ranking.

Latency, VRAM, and the open-weight catch

Cloudflare's 43-benchmark latency table puts median response at 209.3 ms for Clef, 38.8 ms for Clef-flash, and 524.1 ms for Jev. That is about 2.5× faster than Jev at the median for Clef, and about 13.5× for Flash (524.1 ÷ 38.8). Edge hosting on Workers AI is meant to keep the network hop short so you can put the model in front of tool calls: should this ticket escalate, which team owns it, is this crawler allowed.

Open weights under Apache 2.0 are the other half of the story. Michelle Chen, Cloudflare AI Platform group product manager, told The Register that Clef-flash needs about 41 GB of VRAM and Clef about 85 GB at single concurrency with a 64K window. Training datasets stay private, so "open source" here means runnable weights and code, not a full reproducible train. If you already pay for Jev and only need text triage, the 5.7× hosted premium is a hard sell. If you need image-aware decisions, a Jev-compatible swap, or local inference on fat GPUs, Clef is the first serious open alternative in this niche.

Geeknewz verdict

Geeknewz's view: try Clef-flash on Workers AI first if you already sit on Cloudflare and your agent makes dozens of routing checks per session. The $0.09/M sticker and ~39 ms median claim are the practical entry, and the API swap is designed to be a model-name change. Keep full Clef for vision or harder multi-question schemas, and keep Jev if unit economics beat everything else and text is enough. Treat the leaderboard tables as vendor scorecards until independent Decision Index ranks catch up, and budget real GPU memory before you promise a self-hosted Clef box to your team.

Source: Cloudflare Blog; model pages for Clef and Clef-flash; changelog Workers AI post; additional reporting from The Register.