Runtime

llama.cpp 0.6.0 adds Hub downloads and expands current-model support

The runtime release combines a Hugging Face Hub download pipeline with new batch APIs, decision-model serving and explicit support for GLM-5.3-Flash, Clef and Qwen4Exp workflows.

Open Model Weights published6 Oct 2026
Primary sourceggml-org
Source published2026-10-05

ggml-org has released llama.cpp v0.6.0, a substantial runtime update that extends both model support and the project's user-facing download path. The release introduces the `llama_batch_ext` extended batch API with `llama_process`, designed for mixed token and embedding inputs plus MTP and deepstack state embeddings. For local-model users, the Web UI now includes a Hugging Face Hub data layer and model-download pipeline, reducing the distance between discovering a compatible repository and loading it into llama.cpp.

The release notes explicitly add support for GLM-5.3-Flash, described there as a 320B hybrid model, and for the Clef decision model across text and vision. Qwen4Exp gains MTP speculative decoding support. The server also ships `/v1/systemone` for decision models. That last capability overlaps conceptually with the earlier decision-model work already covered by Open Weight Intelligence, but v0.6.0 is broader: it packages the endpoint together with new model support, download plumbing, batch APIs and backend improvements in a tagged runtime release.

The update also includes a Metal tensor-API flash-attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan and an update to ggml v0.26.0. For Open Model Weights, these are compatibility facts rather than model facts. Where a release explicitly names supported model families, runtime evidence can be updated; license, weight-format and publisher fields still come from the model's own repository and primary documentation.

OMW REGISTRY WATCH

What this changes for the evidence layer

Compatibility-layer update only. Runtime-support evidence may be updated where the release explicitly names a model or family, but it does not alter model licenses, ownership or the existence of weight artifacts.

PRIMARY SOURCE

ggml-org

This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.

Open source article ↗