Cloudflare Clef: Open Decision Models Built on Qwen3.8 Arrive on Workers AI
On October 1, 2026, Cloudflare launched Clef and Clef-flash, its first in-house trained models, and opened the weights under Apache 2.0. The release is significant because it moves beyond generative text toward schema-bound decision models that return typed probabilities instead of free-form output. Both models are now hosted on Workers AI and can be swapped in for Typesafe’s Jev with minimal code changes.
What Clef Actually Is — Decision Models vs Text Generation
Decision models are a different species from chat LLMs. Rather than producing paragraphs, Clef reads an input state and a set of typed questions, then returns a probability for every allowed answer. The agent gets a structured decision it can act on immediately: route a ticket, block a request, escalate to a human, or score a user submission against a policy rubric.
Cloudflare positions Clef in the System One family pioneered by Typesafe’s Jev. The API is intentionally compatible. You define questions of three types:
- noul: yes/no, returns the probability the answer is yes.
- choice: pick one option from a set you define, returns the chosen option, per-option probabilities, and a confidence value.
- score: rate against an ordered rubric, returns a probability-weighted score and per-level probabilities.
Because there is no free-form output to parse and no reasoning tokens to wait for, latency drops sharply for hot-path tasks. Cloudflare reports median latencies of 209.3 ms for Clef and 38.8 ms for Clef-flash on its internal benchmarks, versus 524.1 ms for Jev at the median.
System One API Compatibility
The drop-in compatibility means existing Jev integrations can switch by changing the endpoint and model name. Workers AI exposes @cf/cloudflare/clef and @cf/cloudflare/clef-flash. The models accept up to 64 questions per request and a 64K token context window. Clef also includes a vision encoder, allowing up to four images alongside the state, a feature Jev lacks today.
Architecture — Qwen Backbone with Schema Scoring
Cloudflare did not train a new LLM from scratch. Clef freezes Qwen3.8-27B for the full model and Qwen3.5-9B for Clef-flash, then adds a routing head and rank-256 low-rank adapters. During inference, the Qwen backbone runs a prefill-only pass, then scores the valid schema choices in parallel. No autoregressive tokens are generated.
This non-autoregressive scoring is the core innovation. The frozen backbone provides representations, the joint schema head routes evidence to each question, and every option is scored jointly. Training used label-smoothed cross-entropy and a Brier loss for calibration, with RLCD as a secondary objective.
Prefill-Only Inference and LoRA
Because the backbone is frozen, compute is dominated by the initial prefill. The small adapter and routing head keep the parameter delta tiny. Cloudflare says this yields 2.5x faster median latency than Jev for Clef and 13x faster for Clef-flash, with the additional benefit of running on GPUs close to users in Cloudflare’s edge network.
The weights are fully open on Hugging Face under Apache 2.0: Cloudflare/clef and Cloudflare/clef-flash. That means you can run the models locally without paying Workers AI token costs, though you lose the edge placement.
Performance and Benchmarks vs Jev
Cloudflare published results across 10 decision benchmarks. Clef leads on 7 of them. Highlights include BFCL case exact at 98.47 for Clef and 98.76 for Clef-flash versus 95.75 for Jev. BANKING77 macro-F1 is 94.20 for Clef versus 79.74 for Jev. CLINC150+OOS macro-F1 reaches 97.43 for Clef.
On Typesafe’s own workflow evaluations, Clef beats Jev in three of four areas: invoice processing, customer service, and security incidents. Agent trace observability is slightly behind. The models also show strong latency at p95: 238.6 ms for Clef and 122.4 ms for Clef-flash versus 536.0 ms for Jev.
Pricing on Workers AI is $0.24 per million input tokens for Clef and $0.09 for Clef-flash. That is nearly six times the price of Jev at $0.042/M, but the latency and structured output trade-off is the selling point for hot-path agents.
Running Clef Locally — Hardware Reality Check
Local deployment is realistic for evaluation, but the 27B model is demanding. Qwen3.8-27B in Q5_K_M runs around 25 tok/s on an RTX 3090 for generation, but Clef’s prefill-only path changes the profile. You still need VRAM for the full weights plus adapters. On an RTX 4090 or RTX 5090 with 24-32GB VRAM, inference is viable for prototyping. The 9B variant fits comfortably on a 16GB card.
For a dedicated local AI workstation, a high-end GPU with large VRAM, fast system RAM, and NVMe storage matters more than CPU core count. The RTX 5090 class cards with 32GB GDDR7 provide headroom for larger models and batched scoring. A prebuilt with balanced cooling and power delivery keeps sustained prefill performance stable.
What to Watch For
Quantization affects calibration. The Brier loss was tuned for full precision; aggressive quantization can shift probability distributions. If you run locally, validate probabilities on your own decision sets before deploying to production. Also note that vision input requires a vision encoder that is included in the released weights, but image preprocessing must match the training pipeline.
What This Means for Agentic Workflows
Decision models fit the orchestration layer. Use Clef for triage and guardrails, then hand off to a generative LLM for action. Examples Cloudflare highlights include support triage with urgency and team routing, threat intelligence classification with Browser Run, trust and safety scoring against custom rubrics, and agent self-checks before tool calls.
The 64K context window lets you feed larger state snapshots without truncation. Combined with Workers AI’s edge placement, the network round trip stays short. For organizations with strict data residency, the Apache 2.0 weights allow on-prem deployment with no telemetry.
Cloudflare also announced a reinforcement learning fine-tuning service for Clef, starting as a design partner program. That suggests the company expects customers to adapt the routing head to domain-specific schemas, which is cheaper and faster than fine-tuning a full LLM.
Gear for Local Decision Model Testing
If you want to experiment with Clef locally, the hardware that handles Qwen3.8-27B comfortably is the same hardware that handles other 27B-class models. A 32GB VRAM GPU gives you room for full-precision weights and adapters.
The ASUS ROG Astral RTX 5090 White OC Edition 32GB GDDR7 is the current flagship AIB with a 3.8-slot cooler and 4-fan design. It’s built for sustained workloads like local LLM prefill and 4K video editing.
For a turnkey workstation, the Corsair Vengeance i5200 RTX 5090 with Intel Core Ultra 9 285K, 64GB DDR5 and 2+2TB NVMe pairs the GPU with a 24-core CPU and Corsair’s own cooling stack. It’s listed on PC Guide deals as a $500-off premium build.
A high-end prebuilt with AMD is the HP OMEN MAX 45L RTX 5090 with Ryzen 9 9900X3D, 32GB DDR5 and 2TB SSD, which PC Guide highlighted for its 128MB L3 cache and Wi-Fi 7 connectivity. The 32GB VRAM GPU is the key spec for running 27B models locally.
What This Means
Cloudflare’s Clef release legitimizes open-weight decision models as a first-class category. With Apache 2.0 weights built on Qwen3.8-27B and Qwen3.5-9B, the barrier to trying schema-bound inference is low. The performance numbers suggest real latency gains for routing and guardrail tasks, especially when hosted at the edge. For local experimentation, a 24-32GB VRAM GPU remains the practical entry point, and the hardware above matches that requirement.

