WayToClawEarn
Medium impactHugging Face 博客

llama.cpp Adds Native Decision Model Support: 5 Open GGUFs, How to Cut Agent Routing Cost?

llama.cpp PR #29818 adds the /v1/systemone endpoint for native TypeSafe Jev-style decision models: no text generation, returns option probabilities in a single forward pass. Five open GGUF models from 144M to 27B launch first, with the smallest responding in 3ms. Apache 2.0 licensed, supporting choice/score/noul question types.

Edisen Lu · WayToClawEarnVia Hugging Face 博客Published Oct 4, 2026

Reviewed from public sources · AI-assisted drafting under editorial oversight. How we work · Original source

Public-source compilation

Synthesized from public posts/docs. Prefer the original source for primary claims.

How we review content · Primary source · Hugging Face 博客

TL;DR

On October 2, 2026, llama.cpp merged PR #29818 adding native decision model support via the /v1/systemone endpoint, compatible with TypeSafe Jev's System One format. Existing Jev clients need only a new base URL. Five open GGUF models from 144M to 27B launched simultaneously, with the smallest responding in 3ms. This is the key step bringing decision models from cloud APIs to local inference.

What Is a Decision Model

A decision model doesn't generate text. A chat model spends one forward pass per output token, and its output still has to be parsed. A decision model reads the input once, and its answer is always one of your options, with a probability attached.

Typical uses:

  • Routing: Dispatch customer requests to the correct team or tool
  • Moderation: Determine if content violates policies
  • Verification: Check whether an agent's step succeeded
  • Decision-making: Choose an agent's next action

The core advantages are deterministic output and ultra-low latency — no parsing LLM natural language responses, no format errors, structured probability distributions every call.

Five Open Models

llama.cpp's first batch of supported decision model GGUFs:

ModelSizeBase ModelLanguagesImagesLicenseSpeed*
Julia-1144MmmBERT-small50+ languagesNoApache 2.03 ms
Laya421MModernBERT-largeEnglishNoApache 2.05 ms
Kev-4B4BQwen3.5-4B-BaseEnglishNoApache 2.012 ms
lev4BQwen3.5-4BEnglishNoApache 2.036 ms
OpenJev27BQwen3.8-27BEN/DE/FR/HI/ZH/JAYesCC BY-NC 4.043 ms

*Median time to answer one question on a single NVIDIA RTX PRO 6000.

Important notes:

  • Julia-1 and Laya are pure encoder models (mmBERT/ModernBERT) — they don't generate tokens, explaining their speed
  • Kev-4B and lev both use Qwen3.5-4B as base but differ in training; lev supports reasoning chains
  • OpenJev is the only multilingual and image-capable model, but CC BY-NC 4.0 restricts commercial use
  • All except OpenJev are Apache 2.0 licensed, freely usable commercially

Three Question Types

The /v1/systemone endpoint supports three question types:

choice: Provide options with optional descriptions. Returns the top option plus a probability per option.

score: Provide 2-10 levels (lowest first). Returns the expected level (can fall between two levels).

noul: A yes/no question. Returns the probability of "yes."

Quick Start

Start a server:

terminal
llama serve -hf ggml-org/Kev-4B-GGUF

Send a request (asking three questions at once):

terminal
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Customer message: I was charged twice for my order last week and nobody has replied.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "payments, charges, refunds, invoices",
          "shipping": "delivery, tracking, lost or late parcels",
          "technical": "bugs, errors, login problems"
        }
      },
      "angry": {
        "type": "noul",
        "instructions": "Is the customer angry?"
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["can wait", "this week", "today", "right now"]
      }
    }
  }'

Response:

json
{
  "model": "ggml-org/Kev-4B-GGUF",
  "answers": {
    "route": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {"billing": 0.9049, "shipping": 0.0275, "technical": 0.0676},
      "confidence": 0.8574
    },
    "angry": {
      "type": "noul",
      "noul": 0.8208
    },
    "urgency": {
      "type": "score",
      "score": 2.2821,
      "legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
      "probabilities": {"0": 0.036, "1": 0.1937, "2": 0.2225, "3": 0.5478},
      "confidence": 0.2821
    }
  },
  "usage": {"input_tokens": 130, "output_tokens": 0}
}

Note output_tokens: 0 — the decision model generates zero tokens, performing only a single forward pass for scoring.

Image Input

OpenJev (27B) supports image input for document classification, screenshot moderation, etc.:

python
import base64, requests

with open("document.png", "rb") as f:
    image = "data:image/png;base64," + base64.b64encode(f.read()).decode()

response = requests.post("http://localhost:8080/v1/systemone", json={
    "state": "A file uploaded by a customer.",
    "images": [image],
    "questions": {
        "kind": {
            "type": "choice",
            "instructions": "What kind of document is this?",
            "criteria": {"invoice": None, "receipt": None, "contract": None, "other": None},
        },
    },
})
print(response.json()["answers"]["kind"]["choice"])  # invoice

The vision projector downloads automatically — no extra configuration needed.

Several Models, One Server

In router mode, models load on demand and you pick one per request:

terminal
llama serve
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{"model": "ggml-org/Julia-1-GGUF:Q8_0", "state": "...", "questions": {...}}'

/v1/models lists all available model IDs. With a single model loaded, the model field is ignored.

Practical Tips

Try different model sizes. Small models are fast but know less; large ones are more accurate but slower. The HF Decision Index compares them all.

Describe your options. Julia-1 routed "I was charged twice" to shipping with bare labels, but to billing (0.99) once each option had a description.

Pick your confidence cutoff per model. A common pattern is to act on confident answers and send the rest to a human. But the right cutoff depends on the model: a vague ticket ("Hi, quick question about my account") scored 0.25 with Julia-1 but 0.80 with Kev-4B. Test on your own examples before choosing.

Batch your questions. They are answered independently, and Kev-4B, lev, and OpenJev process the state only once.

Try different quantizations. Like any GGUF, these models come in several precisions, e.g. ggml-org/Kev-4B-GGUF:Q8_0.

Connection to the Decision Model Ecosystem

This update brings decision models from cloud APIs to local inference. The ecosystem timeline:

  • TypeSafe Jev (Sep 25): First decision model, cloud API only
  • OpenAI Jev clone (Sep 30): Big lab follows suit
  • AWS Strands Decider 2B (Oct 1): Open-source 2B decision model, local-capable but needs Python environment
  • Cloudflare Clef (Oct 1): Open-source 27B/9B decision models, Workers AI hosted
  • llama.cpp support (Oct 2): Unified local inference framework

Developers can now run all System One-compatible decision models in llama.cpp without building separate inference environments for each. The Hugging Face blog explicitly states "Cloudflare's Clef is next," with more models to follow.

Practical Value for Developers

For teams building agent systems, decision models offer an extremely low-cost routing/verification/moderation solution:

  • Cost comparison: Kev-4B responds in 12ms on a local GPU with zero API cost; the same routing task via GPT-6 Sol costs ~$0.72/task
  • Latency comparison: 3-43ms vs typical LLM calls at 500-3000ms
  • Determinism: Structured probability output, no natural language parsing, no format errors
  • Privacy: Fully local, data never leaves your infrastructure

Applicable scenarios: customer service routing, content moderation, agent step verification, tool selection, urgency assessment — any环节 where you need to choose from a finite set of options.

Model collection: https://huggingface.co/collections/ggml-org/decision-models-6abf80cca3c83f127060a769

PR link: https://github.com/ggml-org/llama.cpp/pull/29818

llama.cpp决策模型JevGGUF开源Agent路由Apache-2.0本地推理

View source →

Educational reference only: cases summarize public sources and may use AI-assisted drafting under editorial review. Not financial advice; outcomes are not guaranteed.