llama.cpp Adds Native Decision Model Support: 5 Open GGUFs, How to Cut Agent Routing Cost?
llama.cpp PR #29818 adds the /v1/systemone endpoint for native TypeSafe Jev-style decision models: no text generation, returns option probabilities in a single forward pass. Five open GGUF models from 144M to 27B launch first, with the smallest responding in 3ms. Apache 2.0 licensed, supporting choice/score/noul question types.
Public-source compilation
Synthesized from public posts/docs. Prefer the original source for primary claims.
TL;DR
On October 2, 2026, llama.cpp merged PR #29818 adding native decision model support via the /v1/systemone endpoint, compatible with TypeSafe Jev's System One format. Existing Jev clients need only a new base URL. Five open GGUF models from 144M to 27B launched simultaneously, with the smallest responding in 3ms. This is the key step bringing decision models from cloud APIs to local inference.
What Is a Decision Model
A decision model doesn't generate text. A chat model spends one forward pass per output token, and its output still has to be parsed. A decision model reads the input once, and its answer is always one of your options, with a probability attached.
Typical uses:
- Routing: Dispatch customer requests to the correct team or tool
- Moderation: Determine if content violates policies
- Verification: Check whether an agent's step succeeded
- Decision-making: Choose an agent's next action
The core advantages are deterministic output and ultra-low latency — no parsing LLM natural language responses, no format errors, structured probability distributions every call.
Five Open Models
llama.cpp's first batch of supported decision model GGUFs:
| Model | Size | Base Model | Languages | Images | License | Speed* |
|---|---|---|---|---|---|---|
| Julia-1 | 144M | mmBERT-small | 50+ languages | No | Apache 2.0 | 3 ms |
| Laya | 421M | ModernBERT-large | English | No | Apache 2.0 | 5 ms |
| Kev-4B | 4B | Qwen3.5-4B-Base | English | No | Apache 2.0 | 12 ms |
| lev | 4B | Qwen3.5-4B | English | No | Apache 2.0 | 36 ms |
| OpenJev | 27B | Qwen3.8-27B | EN/DE/FR/HI/ZH/JA | Yes | CC BY-NC 4.0 | 43 ms |
*Median time to answer one question on a single NVIDIA RTX PRO 6000.
Important notes:
- Julia-1 and Laya are pure encoder models (mmBERT/ModernBERT) — they don't generate tokens, explaining their speed
- Kev-4B and lev both use Qwen3.5-4B as base but differ in training; lev supports reasoning chains
- OpenJev is the only multilingual and image-capable model, but CC BY-NC 4.0 restricts commercial use
- All except OpenJev are Apache 2.0 licensed, freely usable commercially
Three Question Types
The /v1/systemone endpoint supports three question types:
choice: Provide options with optional descriptions. Returns the top option plus a probability per option.
score: Provide 2-10 levels (lowest first). Returns the expected level (can fall between two levels).
noul: A yes/no question. Returns the probability of "yes."
Quick Start
Start a server:
llama serve -hf ggml-org/Kev-4B-GGUFSend a request (asking three questions at once):
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": "Customer message: I was charged twice for my order last week and nobody has replied.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "payments, charges, refunds, invoices",
"shipping": "delivery, tracking, lost or late parcels",
"technical": "bugs, errors, login problems"
}
},
"angry": {
"type": "noul",
"instructions": "Is the customer angry?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["can wait", "this week", "today", "right now"]
}
}
}'Response:
{
"model": "ggml-org/Kev-4B-GGUF",
"answers": {
"route": {
"type": "choice",
"choice": "billing",
"probabilities": {"billing": 0.9049, "shipping": 0.0275, "technical": 0.0676},
"confidence": 0.8574
},
"angry": {
"type": "noul",
"noul": 0.8208
},
"urgency": {
"type": "score",
"score": 2.2821,
"legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
"probabilities": {"0": 0.036, "1": 0.1937, "2": 0.2225, "3": 0.5478},
"confidence": 0.2821
}
},
"usage": {"input_tokens": 130, "output_tokens": 0}
}Note output_tokens: 0 — the decision model generates zero tokens, performing only a single forward pass for scoring.
Image Input
OpenJev (27B) supports image input for document classification, screenshot moderation, etc.:
import base64, requests
with open("document.png", "rb") as f:
image = "data:image/png;base64," + base64.b64encode(f.read()).decode()
response = requests.post("http://localhost:8080/v1/systemone", json={
"state": "A file uploaded by a customer.",
"images": [image],
"questions": {
"kind": {
"type": "choice",
"instructions": "What kind of document is this?",
"criteria": {"invoice": None, "receipt": None, "contract": None, "other": None},
},
},
})
print(response.json()["answers"]["kind"]["choice"]) # invoiceThe vision projector downloads automatically — no extra configuration needed.
Several Models, One Server
In router mode, models load on demand and you pick one per request:
llama serve
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{"model": "ggml-org/Julia-1-GGUF:Q8_0", "state": "...", "questions": {...}}'/v1/models lists all available model IDs. With a single model loaded, the model field is ignored.
Practical Tips
Try different model sizes. Small models are fast but know less; large ones are more accurate but slower. The HF Decision Index compares them all.
Describe your options. Julia-1 routed "I was charged twice" to shipping with bare labels, but to billing (0.99) once each option had a description.
Pick your confidence cutoff per model. A common pattern is to act on confident answers and send the rest to a human. But the right cutoff depends on the model: a vague ticket ("Hi, quick question about my account") scored 0.25 with Julia-1 but 0.80 with Kev-4B. Test on your own examples before choosing.
Batch your questions. They are answered independently, and Kev-4B, lev, and OpenJev process the state only once.
Try different quantizations. Like any GGUF, these models come in several precisions, e.g. ggml-org/Kev-4B-GGUF:Q8_0.
Connection to the Decision Model Ecosystem
This update brings decision models from cloud APIs to local inference. The ecosystem timeline:
- TypeSafe Jev (Sep 25): First decision model, cloud API only
- OpenAI Jev clone (Sep 30): Big lab follows suit
- AWS Strands Decider 2B (Oct 1): Open-source 2B decision model, local-capable but needs Python environment
- Cloudflare Clef (Oct 1): Open-source 27B/9B decision models, Workers AI hosted
- llama.cpp support (Oct 2): Unified local inference framework
Developers can now run all System One-compatible decision models in llama.cpp without building separate inference environments for each. The Hugging Face blog explicitly states "Cloudflare's Clef is next," with more models to follow.
Practical Value for Developers
For teams building agent systems, decision models offer an extremely low-cost routing/verification/moderation solution:
- Cost comparison: Kev-4B responds in 12ms on a local GPU with zero API cost; the same routing task via GPT-6 Sol costs ~$0.72/task
- Latency comparison: 3-43ms vs typical LLM calls at 500-3000ms
- Determinism: Structured probability output, no natural language parsing, no format errors
- Privacy: Fully local, data never leaves your infrastructure
Applicable scenarios: customer service routing, content moderation, agent step verification, tool selection, urgency assessment — any环节 where you need to choose from a finite set of options.
Model collection: https://huggingface.co/collections/ggml-org/decision-models-6abf80cca3c83f127060a769
Topic hub
AI Agent Tutorials & Workflow Guides
Evergreen how-tos for coding agents, content pipelines, and n8n automation—linked to news context and real earn cases.
Explore AI Agent Tutorials & Workflow Guides →Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
DeepSeek + Claude Code Micro SaaS
Run multiple small products on cheap inference
Claude Code bug bounty
Productize agent skills into security services
Related tutorials
Related news
- Xiaomi MiMo-V2.6-Pro Hits Agent Arena: #5 Open-Source, #2 in Confirmed Success — How to Choose Agent Models?
- LangChain Open-Sources Model Router: 64% Cost Cut in Agent Coding with No Quality Loss — How to Build Yours
- Karpathy Proposes 4-Rung LLM Output Ladder: STE100, Diagrams, HTML, Video — How Humans Understand Autonomous Agent Work?
- Comfy Org Launches Comfy Agent: Autonomous Workflow Building on the Canvas, How to Automate ComfyUI Pipelines?