Runtime

vLLM 0.31 targets DeepSeek V4.1 performance and faster model restarts

The October 5 serving release adds a GPU-resident weight-cache daemon, experimental initialized-engine snapshots, new speculative-decoding paths and a broad set of DeepSeek V4.1 Flash optimizations.

Open Model Weights published7 Oct 2026
Primary sourcevLLM
Source published2026-10-05

vLLM 0.31.0 is a substantial serving-layer release focused on restart time, speculative decoding and current large-model architectures. The project says the October 5 release contains 717 commits from 307 contributors. One of the most operationally significant additions is `vllm preload`: a weight-cache daemon designed to keep post-quantized weights resident in GPU memory across engine restarts. The release also adds data-parallel support, MTP draft-model handling, health/readiness plumbing and an experimental CRIU-based snapshot path for restoring a fully initialized TP1 engine.

DeepSeek V4.1 Flash receives a long list of architecture-specific optimizations. The release notes identify FlashMLA mega attention with the model’s NVFP4 compressed KV cache as the SM100 default, alongside sparse MQA indexer work, fused expert-selection paths, tensor-parallel communication fusion, Engram sharding and vision-encoder CUDA graphs. These are implementation claims from the vLLM release and should be treated as runtime evidence rather than independent performance measurements.

Speculative decoding also expands. Model Runner V2 gains draft-model speculative decoding, the release adds a LiLiCorr drafter, DFlash scheduling work, DSpark adaptive verification for Gemma4 and variable-length decode for Kimi-K3. For Open Model Weights, v0.31.0 is important because it creates new compatibility signals for named model families and serving configurations. It does not, on its own, verify model ownership, license terms or weight files; those remain tied to each model’s primary repository.

OMW REGISTRY WATCH

What this changes for the evidence layer

Compatibility-layer update. Explicitly named model support and runtime features may be recorded as source-derived runtime evidence; performance statements remain release-note claims and do not change canonical model specifications, licenses or weight availability.

PRIMARY SOURCE

vLLM

This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.

Open source article ↗