New in llama.cpp: Decision Models

Источник: wavymulder

New in llama.cpp: Decision Models

Source: wavymulder
•Updated: October 3, 2026

llama.cpp server now supports decision models through the /v1/systemone endpoint. You send a state (text, JSON, a screenshot) and typed questions. The model returns a probability for each option in a single forward pass.

The API follows the System One format introduced with TypeSafe's Jev model, so existing clients only need a new base URL. Implementation details are in PR #29818.

What is a decision model? A model that answers by scoring the options you give it, instead of generating text. A chat model spends one forward pass per output token, and its output still has to be parsed. A decision model reads the input once, and its answer is always one of your options, with a probability attached. Typical uses are routing a request, moderating content, checking that an agent's step worked, or choosing its next action.

What is a decision model? A model that answers by scoring the options you give it, instead of generating text. A chat model spends one forward pass per output token, and its output still has to be parsed. A decision model reads the input once, and its answer is always one of your options, with a probability attached. Typical uses are routing a request, moderating content, checking that an agent's step worked, or choosing its next action.

Supported models

*Median time to answer one question, on one NVIDIA RTX PRO 6000.

Find these models in the Decision models collection, with more coming. The community Decision Index shows how they compare.

Quick start

Get the latest llama.cpp from llama.app (or run llama update), then start a model:

A request contains a state and one or more questions. There are three question types:

Send a request with your state and questions:

Response (values rounded):

The full reference is in the server docs.

Images

Some models (OpenJev at the time of writing) can also read images, such as documents or screenshots. The vision projector downloads automatically:

For example, to classify an uploaded document:

The state can also be a list of chat messages. Any image_url part (data URL) is read as an image, same as chat completions.

Several models, one server

In router mode, models load on demand and you pick one per request:

/v1/models lists the ids. With a single model loaded, the model field is ignored.

Tips

  • Try several models, of different sizes. Small models are faster, large ones know more. The Decision Index compares them.
  • Describe your options. Julia-1 routed "I was charged twice" to shipping with bare labels, and to billing (0.99) once each option had a description.
  • Pick your confidence cutoff per model. A common pattern is to act on confident answers and send the rest to a human. The right cutoff depends on the model: a vague ticket ("Hi, quick question about my account") scored 0.25 with Julia-1 but 0.80 with Kev-4B. Test on your own examples before choosing it.
  • Batch your questions. They are answered independently, and Kev-4B, lev and OpenJev process the state only once.
  • Try different quantizations. Like any GGUF, these models come in several precisions, for example llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0.

What's next

New open decision models come out every week, and we'll keep adding the best ones. Cloudflare's Clef is next. Is there one you particularly want? Tell us in the comments.

What this article says

Something is unclear? Ask about the article — I will explain in plain words.

Do not want to dig deeper? We will sort it out for you.