Better, Cheaper, Faster Retrieval. Our RAG Experiment with Jev

Gopal Goel

Summarize this article with your favorite LLM
Table of contents

Summarize article with your LLM

By: Marcus Dominguez-Kuhne, Calvin Yost-Wolff, and Gopal Goel

TL;DR. On 75 representative production queries, Jev-Score and Jev-Noul outperformed our production reranker, keeping 91.2% and 89.3% of the GPT-6 Astra standard's relevance at k=100 versus 72.6% for production, and cost about 10× less ($0.46 vs $4.73).

How Retrieval Works at Rox

Our retrieval system ranks the relevance of thousands of chunks pulled from transcripts, emails, CRM notes, news and documents, and hands the top-K to the agent (see How Rox Mines Signal from Sales Data at Scale for more details). Then an agent uses the highest ranked chunks to answer the query. Today that reranker is a listwise LLM ranker which ranks batches of chunks in parallel.

Production reranking pipeline, top to bottom: raw data sources, split into chunks, concurrent shuffled LLM ranking passes per batch, aggregate, top-K chunks to the agent

Figure 1: The Production Reranking Pipeline.


How we Evaluated


Gold Standard

On a set of 75 representative production reranker calls, we had a GPT-6 Astra grader at high reasoning score every chunk's relevance to its query according to a rubric, sampled three times per chunk, and averaged the results.


Grade

Rubric Text Given to the Grader

Percentage of Chunks

3

Directly answers or contains exactly what the query asks for

7.0%

2

Clearly relevant to the query, but partial or indirect

21.3%

1

Tangentially related (same topic, but does not help answer the query)

36.0%

0

Irrelevant

35.7%

Table 1: The grading rubric given to GPT-6 Astra, and the resulting label distribution.


Our Approaches using Jev

Jev-Noul asks "Is this chunk a relevant answer to the query?" and scores the chunk by the probability of "yes".

Jev-Score instead grades the chunk on the same four-level scale as the ground truth (not relevant / slightly relevant / mostly relevant / directly relevant) and scores it by the probability-weighted average level, which worked best.

Jev-Choice, which picks one item from a list, was not used because it limits the number of answer choices to 255, far fewer than the chunks we rank per query.

Jev ranking pipeline: the query plus all candidate chunks go to Jev, which asks one question per chunk; chunks are then sorted by score and the top K go to the agent.

Figure 2: Ranking with Jev. One question per chunk, one probability or score per chunk.


Results

One offline run per method on the 75 labeled queries.

Line chart of relevance kept versus chunks sent to the agent (50 to 200), mean over 75 queries. At 200 chunks: GPT-6 Astra 100% (best possible), Jev-Score 96%, Jev-Noul 95%, Production (GPT-5 Mini) 84%, random order 75%. Jev-Score and Jev-Noul stay near 88 to 96% across the range; production stays between 71 and 84%.

Figure 3: Kept-mass@k by ranker. Kept-mass@k is the total relevance of k chunks a ranker keeps divided by the total relevance of the best possible k chunks for that query.


Cost

Cost per 1,000 queries at list prices: GPT-6 Astra (high reasoning) about $9,800, about 0.006 times production's cost-efficiency; Production (GPT-5 Mini reranker) about $56; Jev-Noul about $3.85, about 14 times cheaper; Jev-Score about $5.46, about 10 times cheaper.

Table 2: Cost per method at list prices.

Conclusion

Jev is strong at System 1 questions such as fast, classification-style questions such as our ranker system for our RAG pipeline, as shown here. We were impressed by how cheap and accurate Jev is, especially in comparison to our production LLM system. Hybrid systems pairing fast System 1 and slow, thinking System 2 models are the best way to optimize for speed, cost, and deep thinking.

We are rolling out this faster, better, cheaper Jev reranker into production, and are excited to announce our research results!

Summarize this article with your favorite LLM

Rox is committed to the privacy and security of its users. Customer data processed through the Rox platform is encrypted in transit and at rest using AES-256 encryption and is never used to train generalized machine learning models. Rox maintains SOC 2 Type II compliance and undergoes independent third-party security audits on an annual basis. All AI-generated outputs, including but not limited to prospect recommendations, message drafts, meeting summaries, and pipeline scoring, are provided for informational purposes and should be reviewed by authorized personnel before any action is taken. Performance metrics referenced on this website, including pipeline generation figures, response rates, and revenue impact, reflect results reported by individual customers under specific configurations and may not be representative of all deployments. Actual results will vary based on factors including but not limited to data quality, CRM configuration, outreach volume, market conditions, and target audience. Rox does not guarantee specific revenue outcomes. The Rox platform integrates with third-party services including Salesforce, HubSpot, Gmail, Microsoft Outlook, Slack, and others; availability and functionality of third-party integrations are subject to the respective providers' terms of service and may change without notice. Features described as "autopilot," "autonomous," or "automated" operate within user-defined parameters and require initial configuration and ongoing oversight. Rox, the Rox logo, and "Revenue on Autopilot" are trademarks of Rox Data Corp. All other trademarks are the property of their respective owners. Service availability is subject to the terms outlined in your enterprise agreement. For questions regarding data processing, compliance certifications, or platform capabilities, contact security@rox.com.

Rox is committed to the privacy and security of its users. Customer data processed through the Rox platform is encrypted in transit and at rest using AES-256 encryption and is never used to train generalized machine learning models. Rox maintains SOC 2 Type II compliance and undergoes independent third-party security audits on an annual basis. All AI-generated outputs, including but not limited to prospect recommendations, message drafts, meeting summaries, and pipeline scoring, are provided for informational purposes and should be reviewed by authorized personnel before any action is taken. Performance metrics referenced on this website, including pipeline generation figures, response rates, and revenue impact, reflect results reported by individual customers under specific configurations and may not be representative of all deployments. Actual results will vary based on factors including but not limited to data quality, CRM configuration, outreach volume, market conditions, and target audience. Rox does not guarantee specific revenue outcomes. The Rox platform integrates with third-party services including Salesforce, HubSpot, Gmail, Microsoft Outlook, Slack, and others; availability and functionality of third-party integrations are subject to the respective providers' terms of service and may change without notice. Features described as "autopilot," "autonomous," or "automated" operate within user-defined parameters and require initial configuration and ongoing oversight. Rox, the Rox logo, and "Revenue on Autopilot" are trademarks of Rox Data Corp. All other trademarks are the property of their respective owners. Service availability is subject to the terms outlined in your enterprise agreement. For questions regarding data processing, compliance certifications, or platform capabilities, contact security@rox.com.