
Asif Razzaq
· 4 min read
Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence
Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a contextual embedding model for RAG pipelines. Each chunk is embedded with the full document in view. The real change is the training signal. The model learns to retrieve the answer along with the context needed to verify it, not one ‘gold passage.’ P
Is it deployable? Yes, as a self-hosted preview. Weights are on Hugging Face under the MIT license. Loading requires transformers>=5.4.0 with trust_remote_code=True. It is not yet on the Perplexity API. The model card notes that weights and interface may change without backward compatibility.
Why the gold passage falls short
RAG systems split long documents into chunks. A chunk often depends on an entity, heading, or definition stated elsewhere. Contextual models address this with late chunking. The document is encoded in one pass, then pooled per chunk.
Training, however, usually marks one gold chunk per query. Every other chunk becomes a negative, including the sentences that make the answer checkable. Perplexity lists 3 more problems. Binary labels give a coarse signal. LLM annotation cost grows linearly with dataset size. Labels are also tied to one chunking strategy.
How the training works
The teacher is Perplexity’s query-aware context compression model. It reads the query and document together and scores every token.
- Chunk relevance: the mean of the top n token scores inside each chunk.
- Soft target: a temperature-scaled softmax over chunks in the positive document. Chunks in other documents get zero.
- Distillation loss: forward KL divergence between teacher and student distributions.
- Document loss: InfoNCE, where a document scores as its best chunk, inspired by ColBERT’s MaxSim.
Each batch samples a random chunking strategy. Chunks are separated by a learned <|chunk_sep|> token and mean-pooled. The teacher runs only during training, so inference adds no latency or storage.
The model starts from an in-house 9B ColBERT retrieval model. A linear projection outputs 2048 dimensions. Matryoshka training also supports 1024 dimensions. Quantization-aware training enables native int8 embeddings. The release is a model soup of several checkpoints. Training used roughly 430 datasets covering over 50 languages, with no ConTEB data.
Interactive explainer
pplx-embed-v2 Contextual Embedding Explainer ctxHow pplx-embed-v2-context retrieves answers and their evidence
An interactive walkthrough of Perplexity’s contextual embedding preview: why isolated chunks fail, how teacher distillation replaces the single gold passage, and what the reported numbers mean.
01Why context 02How it trains 03Reported results 04Storage mathThree lease files share the sentence “Monthly rent is …”. Only one belongs to 5 Park Avenue. Switch modes and run retrieval.
QUERYWhen does 5 Park Avenue’s lease end and what is the current rent? Isolated sentencesContextual (late chunking) Run retrieval Pick a mode and press Run retrieval.Illustrative example modeled on the lease scenario in Perplexity’s post. Scores are for explanation only, not model outputs.
A context compression model acts as teacher. It scores every token for the query. Those scores are pooled per chunk (mean of the top n tokens) and turned into a soft target, instead of a one-hot gold label.
Chunk boundaries: Sentences2-sentence chunks Temperature0.25 Animate 1 Teacher scores tokens2 Top-n mean per chunk3 Softmax target4 Student matches via KLGold-chunk label (one-hot)
Teacher target (soft, from token scores)
Change the boundaries: the same token scores re-aggregate without re-annotation. That is the “flexible chunk boundaries” property Perplexity describes. Token scores here are illustrative.
context-bench (2,099 queries, 38,894 documents, 2,458,072 sentence chunks, exhaustive ranking). Numbers below are as reported by Perplexity at K = 10.
pplx-embed-v2-context-9b-previewvoyage-context-4 (derived from reported gap) ReplayVoyage values are computed as Perplexity’s figure minus the stated gap (14.4 and 5.0 points). Other Voyage metrics appear only in Perplexity’s chart and are not shown here.
Contextual embeddings store one vector per chunk, same as a normal chunk index. Cost depends on vector size. Perplexity reports that 1024-dim int8 (1 KB) slightly exceeds voyage-context-4 at 2048-dim float32 (8 KB) on its chunk-retrieval suite.
Chunks 1024-d2048-d int8float32 0 vector storage (vectors only, not full index) Bytes per vector– vs 2048-d float32– Chunk-size sensitivity (64 to 512 tokens)81.0% to 79.9%Bytes = dimensions x bytes per value. Sensitivity is mean nDCG@10 across 74 MTEB tasks, as reported by Perplexity.
Sources: Perplexity Research · Model card Built by Marktechpostcontext-bench: a new benchmark
context-bench is built and privately held by turbopuffer to limit training contamination. It holds 2,099 queries over 38,894 documents in 21 domains. Sentence chunking yields 2,458,072 chunks. Median target document length is roughly 6,100 tokens. Queries test 12 contextual capabilities, from pronoun resolution to table structure. Perplexity says the model was submitted blind.
Metrics are Document@K, Answer@K, Evidence Recall@K, and All-Evidence@K. Every model is ranked exhaustively against all chunks, so index settings play no role.
Results
- context-bench at K = 10: 45.5% answer recall, 40.6% evidence recall, 31.1% all-evidence recall.
- Document recall: 15.2% at K = 1 and 61.6% at K = 10.
- vs voyage-context-4: ahead by 14.4 points on answer recall and 5.0 on evidence recall at K = 10.
- ConTEB: highest average nDCG@10 among models shown. pplx-embed-context-v1-4B wins NarrativeQA, and Nemotron-3-Embed-8B wins COVID-QA.
- General retrieval: best average on query-to-chunk tasks. Slightly behind voyage-context-4 on query-to-document.
- Storage: 1024-dim int8 (1 KB per vector) slightly beats voyage-context-4 at 2048-dim float32 (8 KB).
- Chunk size: average score moves from 81.0% to 79.9% between 64 and 512 tokens.
Comparison with the closest competitors
Sources: Voyage docs, model cards linked above. Checked September 30, 2026.
Key Takeaways
- A token-level teacher replaces the single gold-chunk label.
- The model retrieves answers plus supporting evidence in one chunk index.
- context-bench Answer@10 is 45.5%, 14.4 points above voyage-context-4.
- Open MIT weights ship today; Perplexity API access is still pending.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
Original source
This story was published by MarkTechPost and written by Asif Razzaq. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on marktechpost.com


