llama.cpp has added support for a class of models that answer by scoring a fixed set of choices rather than generating a text response token by token. The new server endpoint, `/v1/systemone`, accepts a state such as text, JSON or an image together with typed questions, then returns probabilities for the available options. The API follows the System One format introduced with TypeSafe’s Jev model, which is intended to make the same client pattern portable across compatible backends.
This matters because many agent and automation tasks are fundamentally decisions, not conversations. A router may need to choose one tool, a moderation layer may need to classify one policy state, or an agent may need to decide whether the previous step succeeded. A conventional chat model can perform those tasks, but it still generates output that must be parsed and validated. A decision model instead exposes the choice itself as the output contract.
For the open-weight ecosystem, the llama.cpp integration is more important than a single benchmark number. llama.cpp is one of the major local-inference runtimes and sits underneath a broad set of desktop and developer tools. When a new model class becomes a first-class runtime primitive there, it becomes materially easier to test and deploy outside a hosted API. Open Model Weights should treat this as a compatibility signal rather than a quality judgment: support in a runtime is evidence about deployability, while the usefulness of any particular decision model still depends on its own weights, license, task fit and measured behavior.
What this changes for the evidence layer
Compatibility-layer news. This changes what supported runtimes can do; it does not by itself verify a new model record.
ggml-org / Hugging Face
This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.
Open source article ↗