Runtime

llama.cpp adds decision-model inference through a new System One endpoint

The local inference stack can now score typed options directly instead of generating free-form text, opening a different deployment path for routing, moderation and agent control.

Open Model Weights published2 Oct 2026
Primary sourceggml-org / Hugging Face
Source published2 Oct 2026

llama.cpp has added support for a class of models that answer by scoring a fixed set of choices rather than generating a text response token by token. The new server endpoint, `/v1/systemone`, accepts a state such as text, JSON or an image together with typed questions, then returns probabilities for the available options. The API follows the System One format introduced with TypeSafe’s Jev model, which is intended to make the same client pattern portable across compatible backends.

This matters because many agent and automation tasks are fundamentally decisions, not conversations. A router may need to choose one tool, a moderation layer may need to classify one policy state, or an agent may need to decide whether the previous step succeeded. A conventional chat model can perform those tasks, but it still generates output that must be parsed and validated. A decision model instead exposes the choice itself as the output contract.

For the open-weight ecosystem, the llama.cpp integration is more important than a single benchmark number. llama.cpp is one of the major local-inference runtimes and sits underneath a broad set of desktop and developer tools. When a new model class becomes a first-class runtime primitive there, it becomes materially easier to test and deploy outside a hosted API. Open Model Weights should treat this as a compatibility signal rather than a quality judgment: support in a runtime is evidence about deployability, while the usefulness of any particular decision model still depends on its own weights, license, task fit and measured behavior.

OMW REGISTRY WATCH

What this changes for the evidence layer

Compatibility-layer news. This changes what supported runtimes can do; it does not by itself verify a new model record.

PRIMARY SOURCE

ggml-org / Hugging Face

This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.

Open source article ↗