Inference

Index Team publishes official GGUF builds for its 35B-A3B translation model

The official Index-Translate GGUF repository packages llama.cpp conversions from Q2_K through F16 for the 36B-parameter preview model, with Apache-2.0 metadata and local-serving instructions.

Open Model Weights published7 Oct 2026
Primary sourceIndex Team / Hugging Face
Source published2026-10-03

Index Team has published an official GGUF conversion of its Index-Translate-35B-A3B-preview model, adding a first-party local-inference path for the multilingual translation family. The repository describes Index-Translate as covering 150 languages and supporting terminology- and format-constrained translation as well as controlled dubbing and long-document workflows. Hugging Face marks the GGUF repository Apache-2.0, and the model tree links it directly to the non-GGUF preview checkpoint rather than presenting it as a separate base architecture.

The artifact set spans a wide range of static post-training quantizations produced with llama.cpp. Index Team lists Q2_K at 13.25 GB, Q3 variants from 15.55 GB, Q4_K_M at 21.71 GB, Q5 variants around 25 GB, Q6_K at 29.21 GB, Q8_0 at 37.80 GB and a two-shard F16 conversion baseline at 71.07 GB. The repository also contains multimodal projector files for image-input use; text-only translation can use the plain GGUF files. The model card provides direct llama.cpp serving commands and recommends Q4_K_M as a balance point.

Index Team says each quantization tier was checked against the F16 conversion using GPU validation and generation spot checks before the October 3 publication. Those validation statements remain publisher claims. For the registry, the important distinction is artifact lineage: these are official quantized derivatives of the preview model, so exact GGUF files and formats can be tracked without collapsing the quantized repository and its source checkpoint into unrelated model records.

OMW REGISTRY WATCH

What this changes for the evidence layer

Official quantized-artifact candidate. Preserve the relationship to Index-Translate-35B-A3B-preview rather than treating each quantization as an unrelated base model, and verify exact files, lineage and license evidence before adding or updating a canonical record.

PRIMARY SOURCE

Index Team / Hugging Face

This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.

Open source article ↗