Better, Cheaper, Faster Retrieval. Our RAG Experiment with Jev
Gopal Goel

By: Marcus Dominguez-Kuhne, Calvin Yost-Wolff, and Gopal Goel
TL;DR. On 75 representative production queries, Jev-Score and Jev-Noul outperformed our production reranker, keeping 91.2% and 89.3% of the GPT-6 Astra standard's relevance at k=100 versus 72.6% for production, and cost about 10× less ($0.46 vs $4.73).
How Retrieval Works at Rox
Our retrieval system ranks the relevance of thousands of chunks pulled from transcripts, emails, CRM notes, news and documents, and hands the top-K to the agent (see How Rox Mines Signal from Sales Data at Scale for more details). Then an agent uses the highest ranked chunks to answer the query. Today that reranker is a listwise LLM ranker which ranks batches of chunks in parallel.

Figure 1: The Production Reranking Pipeline.
How we Evaluated
Gold Standard
On a set of 75 representative production reranker calls, we had a GPT-6 Astra grader at high reasoning score every chunk's relevance to its query according to a rubric, sampled three times per chunk, and averaged the results.
Grade | Rubric Text Given to the Grader | Percentage of Chunks |
|---|---|---|
3 | Directly answers or contains exactly what the query asks for | 7.0% |
2 | Clearly relevant to the query, but partial or indirect | 21.3% |
1 | Tangentially related (same topic, but does not help answer the query) | 36.0% |
0 | Irrelevant | 35.7% |
Table 1: The grading rubric given to GPT-6 Astra, and the resulting label distribution.
Our Approaches using Jev
Jev-Noul asks "Is this chunk a relevant answer to the query?" and scores the chunk by the probability of "yes".
Jev-Score instead grades the chunk on the same four-level scale as the ground truth (not relevant / slightly relevant / mostly relevant / directly relevant) and scores it by the probability-weighted average level, which worked best.
Jev-Choice, which picks one item from a list, was not used because it limits the number of answer choices to 255, far fewer than the chunks we rank per query.

Figure 2: Ranking with Jev. One question per chunk, one probability or score per chunk.
Results
One offline run per method on the 75 labeled queries.

Figure 3: Kept-mass@k by ranker. Kept-mass@k is the total relevance of k chunks a ranker keeps divided by the total relevance of the best possible k chunks for that query.
Cost

Table 2: Cost per method at list prices.
Conclusion
Jev is strong at System 1 questions such as fast, classification-style questions such as our ranker system for our RAG pipeline, as shown here. We were impressed by how cheap and accurate Jev is, especially in comparison to our production LLM system. Hybrid systems pairing fast System 1 and slow, thinking System 2 models are the best way to optimize for speed, cost, and deep thinking.
We are rolling out this faster, better, cheaper Jev reranker into production, and are excited to announce our research results!
Similar Articles
We build with the best to make sure we exceed the highest standards and deliver real value.

