> The Recall Ceiling: Why LLM Reranking Cannot Fix Sparse Retrieval
A formal DPI-derived bound proves reranking cannot recover items absent from retrieval. Across 8 datasets: 73-91% zero-recall users make even perfect reranking yield near-zero NDCG. Adaptive-K raises the ceiling by 62-105%.
// TABLE_OF_CONTENTS
The Recall Ceiling¶
The Theorem (Data Processing Inequality)¶
Any reranker operating on K candidates cannot produce items outside that set. At typical R@100 = 5% with 10 relevant items per user, max NDCG@10 ≈ 0.09 regardless of reranker quality.
8-Dataset Evidence¶
| Dataset | R@100 | Zero-Recall Users | Best Ceiling Utilization |
|---|---|---|---|
| Beauty | 8.4% | 88% | 10% |
| Movies | 3.6% | 73% | 12% |
| Electronics | 2.8% | 80% | 7% |
| Toys | 13.8% | 91% | 53% (LambdaMART) |
| MovieLens | 18.9% | 46% | 10% |
LambdaMART (2011) achieves best ceiling utilization — not any modern deep learning reranker.
LLM Reranking Result¶
Qwen2.5-7B ceiling utilization: only 25-32%. 68% of potential gain is unrealized even within the constrained recall set.
Counterintuitive Warm User Finding¶
Warm users (more history) have lower retrieval recall. They've consumed popular items, leaving long-tail test items that CF cannot retrieve.
Adaptive-K: The Fix¶
| Dataset | Fixed R@K | Adaptive R@K | Gain |
|---|---|---|---|
| Beauty | 8.2% | 13.9% | +70% |
| Movies | 3.0% | 6.1% | +105% |
| MovieLens | 18.9% | 36.0% | +91% |
Per-user candidate set sizing based on estimated sparsity raises the ceiling itself.
Practical Rule¶
Before investing in LLM reranking: compute catalog density. Dense domains (MovieLens: 1.96%) make reranking worthwhile. Sparse domains (Movies: 0.27%) make even perfect reranking yield negligible gains.