Hugging Face has released Transformers 5.19.0 with changes that matter to both open-weight model support and large-model deployment. The headline model addition is EmbeddingGemma 2, a Google multimodal embedding architecture built on Gemma 4. According to the release notes, it can encode text, images, audio and video into a shared 768-dimensional space and uses Matryoshka Representation Learning so applications can truncate embeddings to 512, 256 or 128 dimensions. The implementation also exposes configurable visual and video token budgets and allows unused vision or audio towers to be disabled at load time.
The release also changes how mixture-of-experts models are parallelized. Transformers now includes an expert-parallel token-dispatch implementation and makes it the default for Qwen3 MoE and Mellum. Hugging Face says this removes the earlier requirement that expert-parallel size match tensor-parallel size. Trainer support has also been adapted for expert parallelism, while the documentation now covers tensor parallelism with PEFT adapters. Separately, MoE models that compute router logits now expose them through a more consistent output path when requested.
Cache handling receives another set of deployment-relevant fixes. Version 5.19.0 repairs quantized-cache behavior, prevents generate from mutating a user-supplied cache configuration and adds per-layer cache configuration for heterogeneous models. For Open Model Weights, these are runtime and architecture-support facts: they can strengthen compatibility evidence where the release explicitly names a family, but they should not alter checkpoint, license or weight-artifact fields without direct repository evidence.
What this changes for the evidence layer
Runtime and architecture-support evidence. The release can support compatibility signals for model families it explicitly names, but it does not by itself verify a checkpoint, weight list, license or publisher field. EmbeddingGemma 2 should only enter the canonical registry after its own repository artifacts and license evidence are checked.
Hugging Face Transformers
This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.
Open source article ↗