Inference

AutoTrust compresses GEV-26B-Decide to an 18 GB NVFP4 checkpoint

The quantized decision-model checkpoint cuts publisher-reported weight download size from 49.5 GB to 18 GB and weight memory to 17.1 GiB while retaining a Gemma 4 backbone and a separate System 1 decision layer.

Open Model Weights published6 Oct 2026
Primary sourceAutoTrust / Hugging Face
Source published2026-10-03

AutoTrust has published an NVFP4 version of GEV-26B-Decide that targets substantially lower memory use for a multimodal decision model. The repository says its 3,840 routed expert MLPs are quantized to NVIDIA’s NVFP4 format while attention, dense MLPs, routers, the vision tower, language-model head, System 1 adapter, decision head and temperatures remain unchanged in BF16. AutoTrust reports that this reduces the downloadable weight payload from 49.5 GB for the BF16 version to 18 GB and lowers vLLM weight memory from 51.1 GiB to 17.1 GiB.

The serving path is built around the same decision-model interface as the BF16 model. AutoTrust documents a `/v1/decide` route alongside ordinary OpenAI-compatible endpoints, allowing the checkpoint to return calibrated probabilities for typed choices while retaining the Gemma 4 backbone for reasoning. The repository also documents computer-use and robot-arm demonstrations dated October 3. Latency and benchmark figures in the card are publisher measurements; they remain attributed rather than being presented as independent Open Model Weights results.

The licensing distinction is especially important for the evidence layer. AutoTrust states Apache-2.0 for its adapter, decision head and calibration files, but the underlying Gemma 4 base model remains under the Gemma 4 terms. A future registry record should therefore preserve those layered terms rather than labeling the complete artifact simply Apache-2.0. The current discovery pipeline already sees the related GEV-26B-Decide family as a candidate, but exact NVFP4 shard and lineage verification is still required before canonical inclusion.

OMW REGISTRY WATCH

What this changes for the evidence layer

Quantized-model candidate. Keep license layers separate: AutoTrust states Apache-2.0 for its adapter, decision head and calibration files, while the Gemma 4 base remains under Gemma terms. Verify the quantized shard list and relation to the BF16 source record before publication.

PRIMARY SOURCE

AutoTrust / Hugging Face

This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.

Open source article ↗