Cohere Introduces an AI-Judged Search Metric to Credit Results Old Benchmarks Miss
RCP-nDCG@10 gives previously unjudged documents a fresh assessment. Cohere’s validation favors the approach, but deliberately emphasizes cases where the metrics disagree.
Cohere’s RCP-nDCG@10 gives search systems credit for useful results missing from existing relevance labels, using calibrated AI judgments to score retrieved documents. That creates a different development target: Cohere says it optimized its fifth-generation Embed and Rerank models for the metric, so rankings under older benchmarks may differ. The human evaluation is promising but limited: it deliberately focused on contests where metrics disagreed, and small score margins were unreliable.
01
In 289 decided contests across 273 queries, reviewers favored RCP-nDCG 70% of the time when it and conventional nDCG picked different winners.
02
The headline comparison—77% agreement with human-selected winners versus 52% for nDCG—comes from a pool deliberately oversampled for metric disagreements.
03
Reviewers judged 28% of documents labeled irrelevant to be useful, while two human graders gave identical zero-to-four scores only 42% of the time.
A search system can find the right document and still get no benchmark credit if nobody previously labeled it relevant. Cohereintroduced RCP-nDCG@10 on September 30 to address that gap. The new evaluation method uses a calibrated AI judge to assess retrieved documents, and Cohere has released supporting research, code and data.
A fresh judgment for every retrieved document
The established metric, normalized discounted cumulative gain, or nDCG, rewards relevant documents appearing high in search results. But it takes relevance from an existing answer key. Those labels typically come from reviewers assessing documents found by earlier systems, rather than every document that might answer a query.
Cohere’s method keeps the emphasis on ranking useful results near the top. It changes how relevance is assigned, combining two signals rather than asking an AI model for a single grade:
A shared rubric: The judge answers five yes/no questions about each document, using the same criteria for every query.
Document preferences: The judge compares documents in groups, estimating how likely each is to beat the others.
Calibration joins those signals. The comparisons determine order; the rubric anchors relevance scores to a common scale across queries. A useful document can therefore earn credit even if the original answer key never included it.
Agreement with human-selected winners
77%RCP-nDCG
Cohere reports 77% agreement in its study, which deliberately oversampled contests where the two metrics disagreed.
52%Conventional nDCG
Conventional nDCG matched human-selected winners in 52% of the same contests, according to Cohere.
The human test—and its uneven sample
Cohere hired 46 annotators to assess top-five results from competing systems across NanoBEIR, BRIGHT and ViDoRe v3. They could not see system names, rankings or the answer key. Three reviewers assessed each contest, with the majority choosing the winner; 289 contests across 273 queries reached a verdict.
Where the metrics chose different winners, humans backed RCP-nDCG 70% of the time. Where the metrics agreed, both matched reviewers 87% of the time. The headline comparison above is therefore a result from this deliberately selected pool, not a general success rate for all searches.
The study also exposed limits to any tidy relevance score. Reviewers found 28% of documents labeled irrelevant useful. Yet two humans grading the same document agreed exactly on a zero-to-four scale only 42% of the time. Small RCP-nDCG margins were shaky too: below about 0.02, agreement with reviewers was no better than a coin toss.
A different target for model development
Cohere says it optimized its fifth-generation Embed and Rerank models against RCP-nDCG during development. It cautions that those models may not look strongest under legacy nDCG scores. The stated goal is to reward useful information older evaluation sets missed, rather than improve a score for its own sake.
Kenneth Enevoldsen, primary maintainer of the Massive Text Embedding Benchmark, expressed interest in bringing RCP-nDCG into MTEB. Cohere’s research, code and data give other teams material to examine the method; his endorsement describes an intended integration, not a completed leaderboard change.
Sources
cohere.comRCP-nDCG@10: A New Approach to Enterprise Retrieval Quality | Cohere
Reader comments
Newest comments first. Replies stay oldest first.