WayToClawEarn
Medium impactCloudflare Blog

Cloudflare Open-Sources Clef Decision Models: Returns Probabilities Not Text, How to Cut AI Agent Costs

Cloudflare released Clef and Clef-flash, the first decision models. Unlike LLMs, Clef returns typed probabilities for fixed questions instead of free-form text. Built on Qwen3.8-27B and Qwen3.5-9B, Apache 2.0 licensed, latency as low as 38.8ms, $0.09-0.24/M tokens on Workers AI.

Edisen Lu · WayToClawEarnVia Cloudflare BlogPublished Oct 2, 2026

Reviewed from public sources · AI-assisted drafting under editorial oversight. How we work · Original source

Public-source compilation

Synthesized from public posts/docs. Prefer the original source for primary claims.

How we review content · Primary source · Cloudflare Blog

TL;DR

On October 1, 2026, Cloudflare released Clef and Clef-flash—the first models trained by its Workers AI team and the first open-source decision models. Unlike LLMs that generate tokens one at a time, decision models read an input state and a schema of typed questions, returning a probability for every allowed answer with no free-form text. Both models are open-weight under Apache 2.0, compatible with TypeSafe AI's Jev API, with weights available on Hugging Face.

Decision Models vs. LLMs: The Core Difference

LLM output is a token sequence that still needs parsing. A decision model only answers a fixed set of questions about an input, returning probabilities directly:

Clef supports 3 question types:

  1. noul (yes/no): Returns the probability of "yes"
  2. choice (pick one): Selects from named options, returning per-option probabilities and a confidence value
  3. score (rating): Rates against an ordered rubric, returning a probability-weighted score

On Workers AI, a single request carries up to 64 questions and up to 4 images.

Model Specifications

FeatureClefClef-flash
Parameters27B9B
BackboneQwen3.8-27BQwen3.5-9B
LicenseApache 2.0Apache 2.0
Context window65,536 tokens65,536 tokens
Median latency209.3ms38.8ms
Hosted price (input)$0.24/M tokens$0.09/M tokens
Image inputYesYes

Compared to Jev (TypeSafe AI's System One model, released September 15): Jev has 524.1ms median latency, 32K context, closed weights, $0.042/M tokens. Clef-flash is 13.5x faster on latency.

Technical Architecture

Two-Stage Inference

  1. Backbone prefix-only prefill: The Qwen backbone runs a single prefill-only forward pass over the state and questions (no token generation)
  2. Joint Schema Head: A small transformer reads the final hidden states, routes evidence to each question, enables cross-attention between fields, and jointly scores all options. Per-question softmax converts logits to probabilities

Training Method

  • Backbone frozen (Qwen3.8-27B / Qwen3.5-9B)
  • Joint optimization of routing head: rank-256 low-rank adapters (LoRA)
  • Loss: label-smoothed cross-entropy + Brier loss for calibration
  • Secondary objective: RLCD (Reinforcement Learning for Calibrated Decisions) gives partial credit to adjacent ordinal choices

Benchmarks: Where Clef Wins and Where It Doesn't

From Cloudflare's 10-benchmark shortlist of the Decision Index 0.2.1 suite, a Clef model scored highest on 7:

BenchmarkClefJev
BANKING77 (macro-F1)94.2079.74
CLINC150+OOS (macro-F1)97.4389.27
Home appliances (case exact)97.73 (Clef-flash)52.27

But Jev maintains clear leads on knowledge-heavy tasks:

BenchmarkJevClef
GPQA Diamond78.348.0
MMLU-Pro82.765.9
BBH92.973.7

On TypeSafe's own workflow evals, Clef led Jev in 3 of 4 areas by small margins: invoice processing (64.7 vs 61.8), customer service (76.3 vs 76.0), security incidents (62.9 vs 61.7). Jev leads agent trace observability (71.6 vs 68.5).

In Cloudflare's internal threat intelligence workflow, Clef classified a domain in 2.2 seconds vs 4.7 seconds for gpt-oss-120b.

Deployment and Fine-Tuning

Deployment Options

  • Workers AI binding (env.AI.run())
  • REST API
  • AI Gateway
  • Self-hosting: model card documents testing on a single H200 with BF16 weights

RL Fine-Tuning Service

Cloudflare also announced a reinforcement learning service for tuning Clef on private data:

  • Starts with Cloudflare's forward-deployed engineers
  • Self-serve platform coming later
  • Pipeline integrates AI Gateway, Workers AI, Containers, and a new Trainer component
  • Apply via the design partner form

What Decision Models Mean for AI Agents

Decision models solve a fundamental problem with LLMs in deterministic tasks—non-deterministic output. For AI agent architecture, this means:

  1. Agent routing: Use a decision model to determine which tool or agent should handle a request—faster and cheaper than calling a full LLM
  2. Tool selection: Verify whether arguments are grounded before calling a tool
  3. Guardrails: Gate LLM actions with a decision model—"is this tool call grounded?"
  4. Cost reduction: Migrate high-frequency classification decisions from LLM calls to decision models, significantly cutting inference costs
  5. Latency optimization: Clef-flash's 38.8ms latency makes real-time agent gating feasible

Competitive Landscape

Clef enters an emerging "System One model" market:

  • Jev (TypeSafe AI, Sep 15): closed-source, $0.042/M tokens
  • Kev-9B (Jared Palmer): Apache 2.0, 9B + 45.4M LoRA
  • Laya (Convai Innovations): Apache 2.0, 421M params
  • Clef/Clef-flash (Cloudflare): Apache 2.0, 27B/9B

Cloudflare's advantages: global edge network deployment (335+ cities), native Workers AI integration, RL fine-tuning service, and a Jev API-compatible migration path (just change the endpoint and model name).

Conclusion

Cloudflare Clef represents a new paradigm for AI agent infrastructure—not every decision needs a full LLM. Decision models handle high-frequency routing and classification tasks with lower latency, lower cost, and higher determinism, letting LLMs focus on what truly requires generation and reasoning. The combination of Apache 2.0 open source, Workers AI integration, and RL fine-tuning service makes Clef a powerful tool for AI agent developers to cut inference costs and improve response speed. Caveat: all benchmark numbers are vendor-reported with no independent replication yet.

CloudflareClefdecision modelopen source AIApache 2.0Workers AIAI agentQwen3.8

View source →

Educational reference only: cases summarize public sources and may use AI-assisted drafting under editorial review. Not financial advice; outcomes are not guaranteed.