Perplexity Releases a Search Model Trained to Retrieve Answers and Supporting Evidence
The publicly available preview changes how document passages are scored during training. Perplexity claims benchmark leadership, but its results also show where competing models remain stronger.
Listen to this story
The audio brief
Story brief
3 key pointsPerplexity’s new nine-billion-parameter model is designed to retrieve not just the passage containing an answer, but supporting passages that help explain or verify it. Its training uses a compression model to score passage relevance, avoiding the usual choice between treating only one answer chunk as relevant and labeling an entire document relevant. Perplexity reports strong average passage-retrieval results, but the preview does not lead every task or whole-document retrieval; its claims of no added retrieval-time latency or index storage are not comparative speed measurements.
- 01
The September 30, 2026, public preview, pplx-embed-v2-context-9b-preview, is available on Hugging Face.
- 02
The model supports 1,024- and 2,048-dimensional embeddings and native int8 representations.
- 03
Training draws on roughly 430 public and in-house query–document datasets covering more than 50 languages.
Finding the passage that answers a question is not always enough to explain or verify it. On September 30, 2026, Perplexity released pplx-embed-v2-context-9b-preview, a search model trained to retrieve supporting material alongside answer-bearing passages. The public preview is available on Hugging Face. Its central change is in training: useful context gets a relevance score rather than being dismissed because it lacks the answer.
One answer passage versus a chain of evidence
Embedding models turn text into numerical representations that search systems can compare. Contextual embeddings represent each passage, or chunk, with its surrounding document in view. That helps when a passage refers to someone introduced earlier, relies on a section heading, or uses a definition located elsewhere.
Perplexity’s research post identifies a mismatch between that document-aware design and common training labels. A single selected answer chunk—the “gold passage”—is treated as relevant, while other chunks become negative examples. Those negatives can include the very material needed to resolve a reference or supply a complementary fact.
Labeling the entire document relevant creates the opposite problem: it does not identify which passages matter. Perplexity’s approach aims between those extremes. It uses a context-compression model as a teacher, scoring individual tokens—small units of text—against a query and combining those judgments into passage-level relevance scores.
The teacher does not join the live search
The token-level labels also loosen the connection between training and fixed passage boundaries. The same teacher predictions can supervise different ways of splitting a document, without labeling it again. During training, Perplexity randomly selects a chunking strategy and teaches the model both which document matches a query and where relevance lies inside it.
Perplexity says the teacher runs only during training. At retrieval time, the released model produces one contextual embedding per chunk, without an added compression or reranking stage. The method therefore adds no teacher-model latency or extra index storage, according to the company. That is a claim about this training approach, not a measured speed advantage over rival models.
The nine-billion-parameter model supports 1,024- and 2,048-dimensional embeddings, plus native int8 representations that store values as eight-bit integers. Perplexity says its training draws on roughly 430 public and in-house query–document datasets covering more than 50 languages. All passage-level supervision comes from the compression teacher, rather than chunk annotations in those datasets.
An average lead is not an across-the-board win
Perplexity claims state-of-the-art results on the public ConTEB benchmark and turbopuffer’s new context-bench. Its ConTEB comparison shows the highest average passage-retrieval score among the models displayed, but not leadership on every task. Perplexity’s earlier contextual model scores higher on NarrativeQA; Nvidia’s Nemotron model leads on COVID-QA.
The domain comparisons are similarly mixed. Perplexity reports the strongest average chunk-retrieval result, leading in legal, technical and health tasks. Voyage Context 4 scores higher in finance and multilingual retrieval, while Nemotron leads in conversation retrieval. For whole-document retrieval, the preview’s average is slightly below Voyage Context 4.
A test built around supporting context
Context-bench tests the release’s particular ambition: distinguishing near-duplicate documents, finding answers that depend on distant context, and retrieving evidence needed to verify them. It contains 2,099 queries over 38,894 long documents across 21 domains. Perplexity says it developed the model without access to this benchmark and submitted it for blind evaluation; it also says no ConTEB data was used in training.
Turbopuffer keeps context-bench private to reduce the risk of training contamination and preserve its value as a test. Unlike public ConTEB, access to this evaluation runs through the benchmark holder: researchers seeking an evaluation must contact turbopuffer.
Sources
- perplexity.aiContextual embedding beyond the gold passage
Reader comments
Newest comments first. Replies stay oldest first.