Introduction
Contextual embeddings capture the meaning of each passage or chunk in the context of the whole document. Their training and evaluation typically assume a single relevant chunk per query, known as the “gold passage.” However, a gold passage is often not enough in practice. It may contain the answer but lack the supporting context needed to understand or verify it, leaving it ambiguous in isolation.
We’re introducing pplx-embed-v2-context-9b-preview, our new contextual embedding model. The model is trained with a novel approach that uses Perplexity’s context compression model as a teacher. We aggregate its token-level predictions into chunk-level relevance scores, teaching the embedding model to retrieve both answer chunks and supporting context rather than a single gold chunk. The model produces one embedding per chunk with no additional inference cost and supports 1024-dimensional and int8 embeddings.
It achieves state-of-the-art results on context-bench, a new benchmark for context-aware retrieval created and privately held by turbopuffer, and on ConTEB, a widely used public benchmark. Our model was developed independently of context-bench and evaluated as a blind submission.
Context-bench consists of 2,099 queries over 38,894 long documents in 21 domains and evaluates three capabilities of contextual embeddings: document disambiguation among near-duplicate documents, retrieval of gold chunks whose meaning depends on distant context, and recall of the supporting evidence required to verify an answer.
A preview of the model is publicly available on Hugging Face. The benchmark is held privately by turbopuffer to reduce the risk of training contamination and preserve its value for measuring progress. To request an evaluation, please contact turbopuffer at contextbench@turbopuffer.com.
Why context matters for retrieval
Retrieval systems typically divide long documents into smaller chunks that can be indexed and searched independently. Chunking is convenient in practice, but it removes the context in which each chunk originally appeared, creating a trade-off between the granularity of information compression and the context used during encoding. A chunk may refer to an entity introduced earlier, inherit its meaning from a section heading, or rely on definitions located elsewhere in the document. Once extracted, it may no longer contain enough information to be matched reliably to a query. Encoding the entire document as a single vector does not resolve this, since one pooled representation cannot faithfully capture every facet of a long document.
Contextual embedding models address this limitation by representing each chunk in the context of its surrounding content, most commonly through late chunking: the document is encoded in a single pass and chunk representations are pooled afterwards, so that every chunk vector is computed with the whole document in view. Prior work, such as ConTEB and our own pplx-embed-context-v1, has shown that contextualized chunk representations substantially improve retrieval over independently encoded chunks, particularly on long documents.
The limits of gold-chunk supervision
Training contextual embedding models requires supervision that identifies which parts of a document are relevant to a given query. A common approach is gold-chunk annotations: for each query, an annotator, typically a large language model (LLM), identifies the chunk of the document that answers it.
Training usually combines two contrastive losses: a cross-document loss that contrasts the gold chunk with chunks from other documents, and a within-document loss that contrasts it with the other chunks in the same document.
This approach has a few limitations:
- Only one chunk is relevant. It treats every other chunk in the document as a negative, including those with useful supporting information. Document-level supervision operates at the other extreme. It identifies the document as relevant without specifying which chunks within it matter. In practice, relevance often lies between these extremes. The answer-bearing chunk may depend on a few supporting chunks that introduce an entity, resolve a reference, or supply a complementary fact, while the rest of the document is irrelevant. The model must therefore learn the right level of selectivity: broader than a single gold chunk, but more focused than the document as a whole.
- Coarse supervision. Because each annotation identifies a single gold chunk rather than capturing degrees of relevance across the document, it provides limited training signals for the amount of compute spent producing it.
- Costly to scale. Every training pair requires an LLM to read the document and select the relevant chunk, so the annotation cost grows linearly with the size of the training set and caps how much and how diverse the training data can be.
- Sensitive to chunk boundaries. The labels are also tied to the chunk boundaries used during annotation. A gold label assigned to a fixed token-size window does not specify which shorter sentences or longer paragraphs are relevant. Changing the chunking strategy, or training across multiple chunkings to improve robustness to how documents are split at indexing time, can therefore require annotating the corpus again.
A novel approach to contextual learning
We address the limitations of gold-chunk supervision by distilling relevance judgments from a query-aware context compression model. This model serves as a teacher: it reads the query and document jointly and assigns each document token a relevance score. Its predictions offer several advantages for training contextual embeddings:
- Flexible chunk boundaries. Token-level relevance scores can be aggregated over any chosen chunk boundaries, allowing the same teacher predictions to supervise different chunking without re-annotation.
- Continuous relevance. Unlike binary gold-chunk labels, the teacher provides continuous scores that capture degrees of relevance that naturally distinguish answer-bearing chunks, supporting context, and irrelevant content.
- Lower annotation cost. The teacher is compact and specialized in relevance scoring, making it far cheaper to supervise at scale than a general-purpose LLM.
We train our model with a weighted sum of a document-level contrastive loss and a chunk-level distillation loss over batches of query-document pairs. For each query, the paired document serves as the positive, while the remaining documents in the same batch serve as negatives.
For each batch, we randomly select a chunking strategy, split the documents accordingly, and insert separator tokens between chunks. We then process each document in a single forward pass and mean-pool the token representations within each chunk to obtain contextual chunk embeddings. Queries are encoded separately, with mean pooling over their token representations. Finally, we compute the cosine similarity between every query embedding and every chunk embedding in the batch.
Loss function
Document-level contrastive loss
While our embedding model only produces chunk-level similarities, we can naturally derive document-level similarities from them. We consider a document d to be as relevant to a query q as its most relevant chunk. This scoring rule is inspired by the MaxSim operation of ColBERT, applied here to chunks instead of tokens.
Specifically, given a query q with embedding q and a document d as an ordered list of chunks c with embeddings c, we define the document-level similarity as:
s(q,d):=maxc∈dq⊤c.s(q,d):=\max_{c\in d}\mathbf{q}^{\top}\mathbf{c}.
Given these document-level similarities, we can formulate a document-level InfoNCE objective over in-batch negatives.
For a query q, let d⁺ denote its relevant document and 𝒟 the documents in the batch. The contrastive loss is as follows, omitting temperatures for readability:
Ldoc(q,d+)=−logexp(s(q,d+))∑d∈Dexp(s(q,d))\mathcal{L}_{\mathrm{doc}}(q,d^{+})=-\log\frac{\exp(s(q,d^{+}))}{\sum_{d\in\mathcal{D}}\exp(s(q,d))}
Chunk-level distillation loss
The contrastive loss teaches the model which document is relevant to a query. The distillation loss teaches it where relevance lies within that document, which is essential for a contextual model.
During training, we pass each positive query–document pair to our context compression model, which assigns a relevance score to every token in the document. As with document-level scoring, we derive each chunk’s relevance from its most relevant contents. To reduce sensitivity to outlier tokens, including those at chunk boundaries, we define chunk-level relevance as the mean of the top n token scores within the chunk.
For each query, we apply a temperature-scaled softmax to the teacher’s chunk-level relevance scores within the positive document to obtain a target distribution. Chunks in the other documents in the batch receive zero target probability. We train the embedding model to match this target by minimizing the forward KL divergence between the target and its predicted distribution over all chunks in the batch:
Ldist(q,d+)=KL (pT∥pS)\mathcal{L}_{\mathrm{dist}}(q,d^{+})=\mathrm{KL}\!\left(p^{\mathrm{T}}\parallel p^{\mathrm{S}}\right)
pT(cj)=exp(rj)∑ck∈d+exp(rk)p^{\mathrm{T}}(c_j)=\frac{\exp(r_j)}{\sum_{c_k\in d^{+}}\exp(r_k)}
pS(c)=exp(q⊤c)∑d∈D∑ck∈dexp(q⊤ck)p^{\mathrm{S}}(c)=\frac{\exp(\mathbf{q}^{\top}\mathbf{c})}{\sum_{d\in\mathcal{D}}\sum_{c_k\in d}\exp(\mathbf{q}^{\top}\mathbf{c}_k)}
Here, the superscripts T and S indicate the teacher and student, and rⱼ is the teacher’s relevance score for chunk cⱼ in the positive document d⁺. The teacher’s target distribution assigns zero probability to chunks in all other documents in the batch.
Unlike a single gold-chunk label, this target captures degrees of relevance. Supporting chunks receive weight according to the teacher’s scores rather than being treated as negatives, teaching the model to retrieve supporting context as well as the answer.
Combined loss
Our combined training objective is a weighted sum of the document-level contrastive loss and the chunk-level distillation loss. We average this weighted loss over the batch of query–document pairs.
Figure 2 illustrates our training approach. The context compressor serves as a teacher only during training. At inference time, the embedding model produces one contextual embedding per chunk, with no additional compression or reranking stage. The approach improves how contextual embeddings are learned without increasing index storage or adding teacher-model latency at retrieval time.
Training data
The training data consists of roughly 430 public and in-house query–document datasets covering over 50 languages.
We construct in-house pair datasets from PII-filtered production data and by synthesizing relevant queries over long documents. For each forward pass, we select one dataset and sample an entire batch from it to increase the difficulty of in-batch negatives and avoid shortcut learning. None of our training datasets carry chunk-level annotations: all within-document supervision comes from the compression model, and no ConTEB data is used for training. Moreover, no context-bench data was available during the model development.
Training setup
The model starts from an in-house, 9B-parameter ColBERT retrieval model and produces 2048-dimensional embeddings through a linear projection layer. We mark chunk boundaries with a learned <|chunk_sep|> token and form each chunk embedding by averaging its token embeddings. Queries are encoded by the same model, with their token embeddings averaged into a single vector.
We use Matryoshka training to support both 1024- and 2048-dimensional embeddings, and quantization-aware training to support . During training, we gradually increase the learning rate, hold it constant, then reduce it.
The released model is a model soup that averages the weights of several checkpoints from the same training run.
A new benchmark for chunk-level retrieval
Public contextual-retrieval benchmarks such as ConTEB share the assumption behind gold-chunk supervision: each query has a single relevant chunk and a model is rewarded only for ranking that chunk first. They also inherit its limitations.
First, ConTEB tests retrieval on tasks where a chunk is ambiguous without its surrounding context. On these tasks, contextual models outperformed models that encode each chunk independently. However, it does not ask whether a model also surfaces the other parts of the document that make the chunk understandable.
Second, it does not systematically test whether the right document can be found when near-identical documents in the same corpus differ only in the context that surrounds an otherwise identical sentence.
Both capabilities matter when a search agent or user must understand and verify an answer from a few retrieved sentences rather than an entire document. We therefore built context-bench, a controlled benchmark for retrieval over long documents, designed to distinguish the use of document context from general retrieval quality. It tests whether contextual embeddings use information from elsewhere in a document to support three capabilities:
- Document disambiguation: Identify the correct document among near-identical alternatives.
- Answer retrieval: Retrieve answer-bearing chunks whose meaning depends on details elsewhere in the document.
- Evidence retrieval: Retrieve supporting chunks that clarify or substantiate the answer.
Context-bench contains 2,099 queries and 38,894 documents. Sentence length chunks are used, resulting in 2,458,072 total chunks. The queries cover 21 domains including corporate filings, clinical research, travel, entertainment, education, government, legal contracts, and software documentation.
Each query is designed to test one of twelve contextual capabilities. The primary target documents are intentionally long: 1,061 of the 1,197 distinct primary target documents have at least 200 sentences, with a median length of roughly 6,100 tokens. Every query has exactly one correct answer in the corpus. Where the same fact is stated more than once in a gold document, for example in a summary sentence and a table row, any of those sentences are accepted as answers.
An answer route is one valid combination of an answer chunk and the supporting evidence needed to interpret or verify it. Across the 2,099 queries, the primary target-document answer route has one evidence group for 964 queries, two for 735, three or more for 332, and none for 68. Alternative valid answer routes can require different evidence groups.
Table 1: Benchmark statistics.
Statistic
Value
Queries
2,099
Documents
38,894
Sentence chunks
2,458,072
Median primary target document length
6,100 tokens
Domains
Evidence groups per query (primary target-document answer route)
1: 964; 2: 735; 3+: 332; none: 68
Table 2: Twelve contextual capabilities tested by context-bench, with illustrative examples.
Capability
Example query
Answer chunk
Required document context
1. Entity identity
How much did Northlake spend on share repurchases?
The company repurchased $12.8 billion of shares.
The introduction identifies Northlake as “the company.” A similar Southlake report is not relevant.
2. Period or version
Which cartridge fits the M40 revision B?
Install the KC-42 cartridge.
The applicability section says these instructions cover revision B, not revision C.
3. Definitions and aliases
What withdrawal rate did the 15 mg group have?
Zeta had a withdrawal rate of 12.7%.
The report defines Zeta as the 15 mg group. Without that definition, the chunk does not identify the requested group.
4. Table structure
What is the Max model’s battery capacity?
Typical capacity: 4,250 | 4,700 | 5,150.
The table header establishes Mini | Base | Max, measured in mAh. The answer is therefore 5,150 mAh.
5. Conditions and exceptions
How long is battery coverage for commercial use?
Battery coverage lasts 18 months.
A section introduction restricts this coverage to commercial installations; household installations have a different warranty.
6. References and pronouns
When did Mara become deputy mayor?
She took office in 2021.
The preceding narrative establishes that “she” refers to Mara, not Elin, who is also discussed.
7. List order and position
Which song opened the encore?
14. Lantern Road.
A note elsewhere says songs 14–16 constituted the encore. Track 14 is therefore its opening song.
8. Section scope
Which lens was used for the film’s cinematography?
A 35 mm prime was used throughout.
This passage belongs to “Principal Photography,” not the separate section about publicity photographs.
9. Speaker attribution
What reopening date did the mayor promise?
We will reopen on June 12.
The transcript identifies the mayor as the speaker. The same words from a contractor would not establish the mayor’s promise.
10. Changes over time
What notice period applied after the May amendment?
The notice period is now 45 days.
The amendment history establishes that this change took effect in May, replacing the earlier 30-day rule.
11. Causal relationships
Why did the west pump stop?
That fault caused the shutdown.
The investigation identifies “that fault” as a seized bearing in the west pump, and distinguishes it from an unrelated inlet obstruction.
12. Figurative or cross-language meaning
Figurative:
What does “the winter room” symbolize in the poem?
Cross-language:
Which component does the French instruction tell the technician to remove?
Figurative:
It represents exile, rather than a physical room.
Cross-language:
Retirez le capot. (Remove the cover.)
Figurative:
The commentary identifies “it” as the poem’s recurring “winter room” image.
Cross-language:
The guide’s glossary specifies that “capot” means the protective motor hood, not the shipping cover.
Documents and queries
The queries, documents and capabilities being tested are inspired by conversations with turbopuffer’s customers.
To give one example, consider the search needs of a property management firm. These firms have hundreds, sometimes thousands of lease agreements that differ only in the tenants names, dates, and a few other key details. A common query for that dataset might be “when does 5 park avenue’s lease end and what is the current rent?”. Lease agreements can also be lengthy and the critical disambiguating components that would confirm you’ve retrieved the right document (like the date, address, tenant name, etc.) can often be very far away from the key answer sentence containing the monthly lease payment amount. In practice, identical sentences like “Monthly rent is $X,XXX.” can appear across many documents in the corpus.
A good contextual embedding model should use context from across the document to identify the correct document and retrieve both the answer and its supporting sentences. This lets a user or search agent quickly verify the answer and confirm that it comes from the right document.
Answer and evidence labels
For each query, we label “answer” and “evidence” sentences in the target document. A model passes our “reader sufficiency” test if it retrieves an answer sentence together with at least one sentence from each evidence group. Together, these sentences provide enough context for a reader to verify the answer with supporting evidence.
When contextual embedding models do this well, they can dramatically reduce the number of tokens search agents need to review as well as make it much easier for a person to quickly confirm that they have located the correct answer to their original query without having to read the entire document.
Every evidence group must be recovered, but any one of a group’s alternatives is enough to recover it. It’s common to see the same fact stated in several places across these lengthy documents. Retrieving just one sentence from each evidence group meets the reader sufficiency criteria and as such would be accepted as perfect evidence recall.
Metrics
Every document is chunked into sentences, and the same boundaries are used for all models. Two conditions are tested for non-contextual embedding models: one where each sentence is embedded independently and another where the entire document is embedded as a single chunk.
We use the sentence-level condition to measure answer recall and evidence recall. It tests whether a non-contextual model can identify the correct answer and supporting evidence when each sentence is embedded without access to the surrounding document. This is intentionally challenging. In many queries, the answer sentence alone lacks the details needed to distinguish it from similar sentences elsewhere in the corpus.
We use the document-level condition as a document recall baseline. It tests whether fine-grained sentence embeddings provide enough retrieval benefit to justify their additional storage cost relative to storing one vector per document. A sentence-level index creates substantially more vectors than a document-level index, so it should outperform the one-vector-per-document baseline to warrant that cost.
All evaluations use exhaustive rankings against all 2,458,072 chunks. Differences between models therefore reflect the representations, not index configuration. We focus on three key metrics:
Document@K is defined as:
Document@K=1[rank(d∗)≤K]\text{Document@K}=1[\operatorname{rank}(d^{*})\le K]
where d∗d^{*} is the highest-ranked document in DqD_q.
Answer@K is defined as:
Answer@K=1[Aq∩CK(q)≠∅]\text{Answer@K}=1[A_q\cap C_K(q)\ne\varnothing]
where CK(q)C_K(q) is the top KK chunks over the whole corpus and AqA_{q} is the set of accepted answer chunks.
Evidence Recall@K is measured within d∗d^{*} as follows: Rank the sentences of d∗d^{*} by similarity, remove the highest-ranked answer sentence, and let RK(q)R_K(q) be the top KK that remain. For evidence groups Gq={g1,…,gm}G_q=\{g_1,\ldots,g_m\}, group gjg_j is recovered if some alternative E∈gjE\in g_j satisfies E⊆RK(q)E\subseteq R_K(q), and
EvidenceRecall@K=1[rank(d∗)≤10]⋅1m∑j=1m1[gj recovered].\text{EvidenceRecall@K}=1[\operatorname{rank}(d^{*})\le10]\cdot\frac{1}{m}\sum_{j=1}^{m}1[g_j\ \text{recovered}].
Document@K asks whether contextualization lets the right document beat siblings and distractors that share its vocabulary.
Answer@K asks whether the answer sentence itself ranks among the top chunks.
Evidence Recall@K asks how much of the supporting context is recovered among the correct document’s top K chunks, after excluding the highest-ranked answer chunk. Removing the answer keeps evidence recall from double-counting what Answer@K already measures, and leaves all KK slots for supporting context. Requiring the gold document to be in the top 10 retrieved documents means no model receives evidence credit for a document it would not have surfaced. We also report All-Evidence@K, the fraction of queries with every group recovered.
Together, these metrics decompose the challenges contextual-embedding models face: find the right document, find the right answer, and recover the minimal context that shows why the answer is correct.
Figure 3 illustrates the metrics with the property management example.
Evaluation results
We evaluate our model against contextual and non-contextual embedding models on four fronts.
Our contextual baselines are pplx-embed-context-v1-4B and voyage-context-4. Our non-contextual baselines are voyage-4-large and NVIDIA’s Nemotron-3-Embed-8B-BF16, which encode chunks independently. The baseline set varies by experiment, as indicated in each figure.
We start with context-bench, which directly tests the capabilities that motivate this work, and then report results on ConTEB, the standard public benchmark for contextual embeddings. To check that these gains do not come at the expense of general retrieval quality, we also evaluate domain-specific document and chunk retrieval benchmarks that we built from public datasets. Finally, we examine two practical properties for deployment: retrieval quality versus storage costs per vector, and sensitivity to the chunk size used at indexing time. All evaluations encode documents of up to 32,768 tokens in a single pass.











![Scrunch vs. Peec AI: Choosing the right AEO tool [2026]](https://cdn.pixabay.com/photo/2013/04/22/00/50/screw-106361_1280.jpg)
![AEO checker tools that measure answer engine visibility [2026]](https://cdn.pixabay.com/photo/2020/03/09/21/06/distance-4917124_1280.jpg)