Measure retrieval on its own before touching the model
Two hundred questions with the known correct document for each. Recall at 5 is the number. The system's accuracy can never exceed it, no matter which model sits on top.
Ours was 61 percent and the proposal on the table was a bigger model. Six weeks of chunking, hybrid search and a reranker took it to 91 with the same model.
ragtesting
Longer version: the post this came from.