Model release

Perplexity Opens pplx-embed-v2-late Retrieval Models

Perplexity publishes MIT-licensed 0.6B and 9B late-interaction embeddings for text, images and documents, with a shared index-query space and explicit storage tradeoffs.

Open Model Weights published11 Oct 2026
Primary sourcePerplexity official Hugging Face model cards
Source published2026-10-07

What is pplx-embed-v2-late? It is Perplexity's newly published family of open-weight, late-interaction retrieval models for text, images and visual documents. The release comprises separate 0.6B and 9B checkpoints on Hugging Face. Both are marked MIT-licensed and use a shared representation, so an application can build a document index with the larger model and issue queries using the smaller model. This is a model release for search and retrieval, not a new general-purpose chat assistant.

Traditional single-vector retrieval compresses a passage into one embedding. Perplexity's models instead emit a 128-dimensional vector per token and compare query and document vectors with a MaxSim-style late-interaction score. This retains finer-grained matching evidence but can increase index storage and serving complexity compared with one vector per document. The same family covers image inputs and rendered document pages, making it relevant to multimodal search and retrieval-augmented generation pipelines.

The two official model cards describe a Qwen3.5-based architecture with bidirectional attention and explain how the models were distilled from an internal teacher. Perplexity reports 340 million active parameters for the smaller model and 7.4 billion for the larger model, alongside the nominal 0.6B and 9B class names. The card also lists public ViDoRe v3 retrieval scores. Those are publisher-reported evaluations; Open Model Weights has not independently replicated the training, ranking metrics or resource requirements.

For developers, both downloadable checkpoints are concrete artifacts rather than promised weights. The public repositories provide Safetensors and Sentence Transformers integration guidance, including MultiVectorEncoder for late-interaction queries and documents. The larger index and smaller query model should be tested together for quality, latency and vector-storage cost. An available downloadable model does not automatically establish a currently supported managed API endpoint or its price. The OMW registry should verify each repository separately, preserve licensing and exact revisions, and avoid presenting the manufacturer's benchmark numbers as OMW measurements.

OMW REGISTRY WATCH

What this changes for the evidence layer

Two distinct open-weight candidates, not one generic embedding record: verify both Perplexity Hub repositories, their model-specific configs, licenses, Safetensors and revisions. Distinguish nominal model size from active parameters and record the stated shared 128-dimensional token space as publisher documentation. ViDoRe scores must remain attributed, and hosted API availability is not implied by downloadable checkpoints.

PRIMARY SOURCE

Perplexity official Hugging Face model cards

This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.

Open source article ↗