Cloudflare Open-Sources Clef Decision Models: Returns Probabilities Not Text, How to Cut AI Agent Costs
Cloudflare released Clef and Clef-flash, the first decision models. Unlike LLMs, Clef returns typed probabilities for fixed questions instead of free-form text. Built on Qwen3.8-27B and Qwen3.5-9B, Apache 2.0 licensed, latency as low as 38.8ms, $0.09-0.24/M tokens on Workers AI.
Public-source compilation
Synthesized from public posts/docs. Prefer the original source for primary claims.
TL;DR
On October 1, 2026, Cloudflare released Clef and Clef-flash—the first models trained by its Workers AI team and the first open-source decision models. Unlike LLMs that generate tokens one at a time, decision models read an input state and a schema of typed questions, returning a probability for every allowed answer with no free-form text. Both models are open-weight under Apache 2.0, compatible with TypeSafe AI's Jev API, with weights available on Hugging Face.
Decision Models vs. LLMs: The Core Difference
LLM output is a token sequence that still needs parsing. A decision model only answers a fixed set of questions about an input, returning probabilities directly:
Clef supports 3 question types:
noul(yes/no): Returns the probability of "yes"choice(pick one): Selects from named options, returning per-option probabilities and a confidence valuescore(rating): Rates against an ordered rubric, returning a probability-weighted score
On Workers AI, a single request carries up to 64 questions and up to 4 images.
Model Specifications
| Feature | Clef | Clef-flash |
|---|---|---|
| Parameters | 27B | 9B |
| Backbone | Qwen3.8-27B | Qwen3.5-9B |
| License | Apache 2.0 | Apache 2.0 |
| Context window | 65,536 tokens | 65,536 tokens |
| Median latency | 209.3ms | 38.8ms |
| Hosted price (input) | $0.24/M tokens | $0.09/M tokens |
| Image input | Yes | Yes |
Compared to Jev (TypeSafe AI's System One model, released September 15): Jev has 524.1ms median latency, 32K context, closed weights, $0.042/M tokens. Clef-flash is 13.5x faster on latency.
Technical Architecture
Two-Stage Inference
- Backbone prefix-only prefill: The Qwen backbone runs a single prefill-only forward pass over the state and questions (no token generation)
- Joint Schema Head: A small transformer reads the final hidden states, routes evidence to each question, enables cross-attention between fields, and jointly scores all options. Per-question softmax converts logits to probabilities
Training Method
- Backbone frozen (Qwen3.8-27B / Qwen3.5-9B)
- Joint optimization of routing head: rank-256 low-rank adapters (LoRA)
- Loss: label-smoothed cross-entropy + Brier loss for calibration
- Secondary objective: RLCD (Reinforcement Learning for Calibrated Decisions) gives partial credit to adjacent ordinal choices
Benchmarks: Where Clef Wins and Where It Doesn't
From Cloudflare's 10-benchmark shortlist of the Decision Index 0.2.1 suite, a Clef model scored highest on 7:
| Benchmark | Clef | Jev |
|---|---|---|
| BANKING77 (macro-F1) | 94.20 | 79.74 |
| CLINC150+OOS (macro-F1) | 97.43 | 89.27 |
| Home appliances (case exact) | 97.73 (Clef-flash) | 52.27 |
But Jev maintains clear leads on knowledge-heavy tasks:
| Benchmark | Jev | Clef |
|---|---|---|
| GPQA Diamond | 78.3 | 48.0 |
| MMLU-Pro | 82.7 | 65.9 |
| BBH | 92.9 | 73.7 |
On TypeSafe's own workflow evals, Clef led Jev in 3 of 4 areas by small margins: invoice processing (64.7 vs 61.8), customer service (76.3 vs 76.0), security incidents (62.9 vs 61.7). Jev leads agent trace observability (71.6 vs 68.5).
In Cloudflare's internal threat intelligence workflow, Clef classified a domain in 2.2 seconds vs 4.7 seconds for gpt-oss-120b.
Deployment and Fine-Tuning
Deployment Options
- Workers AI binding (
env.AI.run()) - REST API
- AI Gateway
- Self-hosting: model card documents testing on a single H200 with BF16 weights
RL Fine-Tuning Service
Cloudflare also announced a reinforcement learning service for tuning Clef on private data:
- Starts with Cloudflare's forward-deployed engineers
- Self-serve platform coming later
- Pipeline integrates AI Gateway, Workers AI, Containers, and a new Trainer component
- Apply via the design partner form
What Decision Models Mean for AI Agents
Decision models solve a fundamental problem with LLMs in deterministic tasks—non-deterministic output. For AI agent architecture, this means:
- Agent routing: Use a decision model to determine which tool or agent should handle a request—faster and cheaper than calling a full LLM
- Tool selection: Verify whether arguments are grounded before calling a tool
- Guardrails: Gate LLM actions with a decision model—"is this tool call grounded?"
- Cost reduction: Migrate high-frequency classification decisions from LLM calls to decision models, significantly cutting inference costs
- Latency optimization: Clef-flash's 38.8ms latency makes real-time agent gating feasible
Competitive Landscape
Clef enters an emerging "System One model" market:
- Jev (TypeSafe AI, Sep 15): closed-source, $0.042/M tokens
- Kev-9B (Jared Palmer): Apache 2.0, 9B + 45.4M LoRA
- Laya (Convai Innovations): Apache 2.0, 421M params
- Clef/Clef-flash (Cloudflare): Apache 2.0, 27B/9B
Cloudflare's advantages: global edge network deployment (335+ cities), native Workers AI integration, RL fine-tuning service, and a Jev API-compatible migration path (just change the endpoint and model name).
Conclusion
Cloudflare Clef represents a new paradigm for AI agent infrastructure—not every decision needs a full LLM. Decision models handle high-frequency routing and classification tasks with lower latency, lower cost, and higher determinism, letting LLMs focus on what truly requires generation and reasoning. The combination of Apache 2.0 open source, Workers AI integration, and RL fine-tuning service makes Clef a powerful tool for AI agent developers to cut inference costs and improve response speed. Caveat: all benchmark numbers are vendor-reported with no independent replication yet.
Topic hub
AI Agent Tutorials & Workflow Guides
Evergreen how-tos for coding agents, content pipelines, and n8n automation—linked to news context and real earn cases.
Explore AI Agent Tutorials & Workflow Guides →Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
DeepSeek + Claude Code Micro SaaS
Run multiple small products on cheap inference
Claude Code bug bounty
Productize agent skills into security services
Related tutorials
Related news
- DeepSeek Harness v0.2 Desktop Launch: Zero-Config AI Coding Agent, How to Choose?
- ElevenLabs Doubles to $22B Valuation with Eleven v4 Turbo: How to Choose Voice AI Agents?
- OpenAI DevDay 2026: GPT-6.1 Sol at 1/5 Astra Price, Dots Autonomous Agent, Codex Cloud — What Developers Need to Know
- Plugin4Shell: Zero-click RCE hits four major AI coding agents — how to fix