My semantic video search project finds short passages from video transcripts and gives them to a model to answer a question. It currently ranks those passages by one semantic search method and has no labelled test set. That makes every proposed improvement, from keyword search to a reranker or a different chunk size, an untested idea.
This note turns those ideas into a measurement plan: what to measure, how to build a fair test set, how far to trust an automated judge, and which experiments published evidence suggests running first.
How to read the diagrams
- Retrieval: finding the right chunks
- Generation: what the model does with them
- People: labels, judgements, and the questions themselves
- Published evidence that a change helps
- Traps that make an evaluation lie
The system being measured
Retrieval-augmented generation (RAG) first finds source material, then gives it to a model that writes an answer. A transcript is divided into chunks, short passages small enough to retrieve and quote. An embedding represents a chunk or question as numbers so nearby meanings can be found by vector search.
The retriever returns a ranked list of chunks. The generator reads those chunks and writes the answer. BM25 is a keyword-ranking method that rewards query terms while reducing the influence of very common words. Hybrid search combines keyword and vector results. Reciprocal rank fusion (RRF) is one way to do that: it gives a result points from its position in each ranked list, then adds the points. A reranker reads the question and each candidate together, then produces a more expensive relevance score for a smaller candidate set.
Fixed evaluation input
Questions1
human-written + reference timestamps
System under test
Retriever variant2
dense · hybrid · reranked
Ranked chunks3
top 5 · 10 · 20
Judgement
Relevance labels4
0 · 1 · 2 · 3 + correct start
Measurement and decision
Metrics5
recall · MRR · nDCG · time error
Decision6
keep · reject · investigate a slice
Relevance labels say how well a chunk answers one question. This plan uses 0 for irrelevant, 1 for related, 2 for useful but incomplete, and 3 for the exact evidence. An LLM judge is a model asked to assign those labels; human labels remain the reference used to test whether the automated judge is dependable.
Retrieval and generation are different measurements
A RAG answer can be wrong for two unrelated reasons: the right passage was never retrieved, or it was retrieved and the model misused it. Measuring only the final answer mixes the two, so a change that improves retrieval can look like no change at all if the model then ignores the extra context. In Lost in the Middle, performance was often strongest when the needed information appeared near the beginning or end of the context and weaker when it appeared in the middle. The size of that effect depended on the model and task, but the lesson is useful here: better recall does not automatically produce better answers.
So the plan measures the retriever on its own first, with no model involved, and the generator second.
What to measure
- metric
- Recall@k
- what it measures
- Share of the relevant chunks that appear in the top k
- when it's the right one
- When a model will read all k chunks anyway: did the evidence make it in?
- metric
- Hit rate@k
- what it measures
- Share of questions with at least one relevant chunk in the top k (DPR's 'top-20 accuracy')
- when it's the right one
- When one good chunk is enough to answer
- metric
- MRR
- what it measures
- Mean of 1 / rank of the first relevant result (0 if none)
- when it's the right one
- When the first good result is what matters, as with the agent's first search
- metric
- nDCG@k
- what it measures
- Rewards relevant results near the top, and handles graded relevance (perfect vs partly relevant)
- when it's the right one
- The standard summary: BEIR and MTEB both report nDCG@10 as the main number
- metric
- Start-time error
- what it measures
- Distance in seconds between the retrieved chunk's start and the moment the answer begins
- when it's the right one
- Specific to timestamped media: a right chunk that seeks to the wrong moment is half a failure
Here is one worked example. These five chunks and scores are illustrative; they are not results from the project. Suppose the evaluation pool contains exactly three chunks labelled relevant to the question, and all three appear below. A relevance grade of 1, 2 or 3 counts as relevant for recall and MRR. That broad threshold answers a practical retrieval question: did the candidate set contain material that could help the next stage? The stricter grade 2-or-higher result will also be reported as a sensitivity check, because a merely related chunk may not contain enough evidence to answer.
- rank
- 1
- retrieved chunk
- How IVFFlat partitions the vector space
- grade
- 0 · irrelevant
- first relevant?
- no
- rank
- 2
- retrieved chunk
- HNSW uses more memory but gives strong recall and low query latency
- grade
- 3 · exact answer
- first relevant?
- yes
- rank
- 3
- retrieved chunk
- The graph stores several links for each vector
- grade
- 1 · related
- first relevant?
- already found
- rank
- 4
- retrieved chunk
- How cosine distance is computed
- grade
- 0 · irrelevant
- first relevant?
- not applicable
- rank
- 5
- retrieved chunk
- Index builds can take longer as the collection grows
- grade
- 2 · useful
- first relevant?
- already found
Recall@5 asks how many known relevant chunks appeared: 3 retrieved ÷ 3 relevant = 1.00. It ignores order, so putting the exact answer at rank 2 rather than rank 5 makes no difference.
MRR means mean reciprocal rank. For this single question, the first relevant result is rank 2, so its reciprocal rank is 1 ÷ 2 = 0.50. Across a test set, average that value for every question; a question with no relevant result contributes zero.
nDCG@5 keeps the grades and the order. This example uses gain 2^grade − 1 and discount log₂(rank + 1). Its discounted cumulative gain is:
rank 2: (2³ − 1) / log₂(3) = 4.417
rank 3: (2¹ − 1) / log₂(4) = 0.500
rank 5: (2² − 1) / log₂(6) = 1.161
DCG@5 = 6.077The ideal order is grades 3, 2, 1 at ranks 1, 2, 3. Its IDCG@5 is 9.393, so nDCG@5 = 6.077 ÷ 9.393 = 0.647. Normalising by the ideal makes questions with different numbers of relevant chunks comparable on a 0–1 scale.
The last row matters for transcripts. The TREC Podcasts track, which evaluated search over 100,000 podcast episodes, graded segments partly on whether they are "an ideal entry point for a human listener", and earlier speech-retrieval work (CLEF's mGAP measure) penalised the distance between the retrieved start and the true start, up to a limit of 150 seconds (Eskevich et al., 2012). The project's timestamps are the whole point of its answers, so start-time error gets measured alongside ranking quality.
For generation, the useful measurements are the ones RAGAS popularised: faithfulness (the share of the answer's claims that the retrieved context supports), answer relevance, and context relevance. In the RAGAS paper's own validation, the automated scores agreed with human annotators 95% of the time for faithfulness but only 78% for answer relevance and 70% for context relevance, which they called "the hardest quality dimension to evaluate." Those are results for that paper's judge and datasets, not a reliability guarantee. Faithfulness is the first automated signal to test against this project's human labels.
Building a test set that doesn't flatter one retriever
The test set decides the winner more often than the retriever does. Two traps are well documented.
Questions written while looking at the passage favour keyword search. The ORQA authors pointed out that in SQuAD, question writers "are also prompted with a specific piece of evidence for the answer, leading to artificially large lexical overlap between the question and evidence." The DPR paper measured the effect: on SQuAD, BM25 retrieved an answer-bearing passage in the top 20 for 68.8% of questions against DPR's 63.2%, while on Natural Questions, where real users wrote the questions without seeing the answer, the order reversed, 59.1% against 78.4% (Karpukhin et al., 2020). A test set generated by asking a model to write questions about a chunk has the same shape, so it will overstate how well keyword matching works.
Unjudged is not irrelevant. Most test sets label a few passages per question and treat everything else as wrong. BEIR found that 31.8% of TAS-B's top-10 results were unjudged, against 6.4% for BM25, because the original judgement pools came from keyword-based systems. In a separate TREC-COVID reannotation experiment, ANCE's nDCG@10 rose from 0.654 to 0.735 after missing judgements were filled in. Those numbers come from different retrievers and experiments; together they show why each compared system must contribute candidates to the judgement pool.
The plan for this project's test set follows from both:
- Begin with roughly 50 questions, written by people who watched the videos rather than by a model reading chunks. Treat that as a pilot, not a magic minimum. TREC evaluations have often used about 50 topics, while Voorhees and Buckley show that topic-set size changes experiment error. Add questions, or repeat the comparison over resampled question sets, until the uncertainty is small enough for the decision being made. ARES's separate suggestion of about 150 labelled datapoints applies to judge calibration, not to the number of unique questions (Voorhees and Buckley, 2002; Saad-Falcon et al., 2023).
- Graded relevance on four levels, as TREC Deep Learning uses: 3 contains the exact answer, 2 contains it less clearly, 1 is related but doesn't answer, 0 is irrelevant.
- For every question, the correct start time in seconds.
- Pooled judging: every system being compared contributes its top results to the pool that gets labelled, so no retriever is scored against labels that only another retriever could have found.
How far to trust an LLM judge
Labelling every result by hand doesn't scale, so an LLM judge will do most of the labelling, and it has known biases. A widely cited study, Judging LLM-as-a-Judge, found that GPT-4 agreed with human preferences 85% of the time on clear-cut comparisons, slightly more than humans agreed with each other (81%). The same paper found that the order of the two answers changed the verdict: GPT-4 gave the same judgement when the answers were swapped only 65% of the time.
- bias
- Position
- measurement
- Consistent verdict after swapping answer order: GPT-4 65%, GPT-3.5 46.2%, Claude-v1 23.8%
- mitigation used in this plan
- Judge each pair twice in both orders; count a win only if both agree
- bias
- Verbosity
- measurement
- A padded, repetitive answer fooled GPT-3.5 and Claude-v1 in 91.3% of cases, GPT-4 in 8.7%
- mitigation used in this plan
- Grade relevance of a chunk to a question, not preference between answers
- bias
- Self-preference
- measurement
- GPT-4 gave itself a 10% higher win rate, Claude-v1 25% (the authors note this can't prove bias)
- mitigation used in this plan
- Use a different model to judge than the one that generates
- bias
- No reference
- measurement
- On math questions, judge failures dropped from 14 of 20 to 3 of 20 when given a reference answer
- mitigation used in this plan
- Give the judge the human-written reference answer and timestamp
Agreement percentage alone can hide a problem: later work found that judges can agree often while still assigning materially different scores, and that some judges lean toward generous grades (Thakur et al., 2024). Two people will independently label a calibration subset and adjudicate disagreements before it becomes the reference. The first calibration pool will contain at least 150 query-and-chunk judgements, following ARES's input guidance, and will grow if rare grades or question types have too few examples. The model judge will then be compared with that reference using weighted Cohen's kappa, because the four labels are ordered and confusing 3 with 2 is less severe than confusing 3 with 0. Raw agreement, the confusion matrix, and per-grade precision and recall will be shown beside it; one summary coefficient cannot reveal which errors the judge makes.
What to try first, and why
Each experiment below changes one thing, and each is here because published work measured a real effect on a comparable problem.
Anthropic's contextual retrieval: share of queries where the right chunk missed the top 20
Retrieval failure rate
Adding a short model-written context to each chunk, then contextual BM25, then a reranker cut failures by 35%, 49%, and 67%.
y: percent of queries (1 − recall@20)
Microsoft's hybrid retrieval measurements across four customer datasets
Ranking quality
On queries that were mostly exact keywords, vector search scored 11.7 against keyword search's 79.2.
y: NDCG@3
- project gap → experiment
- No evaluation set → build and freeze the first labelled split
- why this experiment
- Every later comparison needs the same questions, labels and metric code
- decision metrics
- judge agreement first; then baseline recall@20, MRR, nDCG@10 and start-time error
- project gap → experiment
- Dense retrieval only → add full-text search and merge with RRF
- why this experiment
- BEIR found BM25 remained a strong out-of-domain baseline; Microsoft's four-dataset study scored hybrid above vector alone (48.4 vs 43.8 nDCG@3)
- decision metrics
- nDCG@10 and recall@20; slice names, jargon and exact quotes
- project gap → experiment
- No reranker → cross-encode the top 50 candidates
- why this experiment
- BEIR reported BM25 plus a cross-encoder above BM25 on 16 of 18 datasets, with an 11% average gain and about 450 ms GPU latency in that setup
- decision metrics
- nDCG@10, MRR and added p50/p95 latency
- project gap → experiment
- Dense retrieval only → add video and section context to each chunk
- why this experiment
- Anthropic reported a 35% reduction in top-20 retrieval failures from contextual embeddings on its evaluated corpora
- decision metrics
- recall@20; compare separately from hybrid so the cause is identifiable
- project gap → experiment
- Chunk start times → fix overlap timestamps, then vary size and overlap
- why this experiment
- Timestamp correctness is a known project bug. Microsoft used 512-token chunks with 25% overlap for all four customer datasets; Chroma found that the best size and overlap depended on the embedding model and that overlap can add redundant tokens
- decision metrics
- start-time error first; then recall@20 and retrieved-token efficiency for each chunking variant
- project gap → experiment
- Captions only → create an ASR and named-entity slice
- why this experiment
- Automatic captions can fail on names, which is also where keyword retrieval may help
- decision metrics
- all retrieval metrics split by caption source and named-entity presence
RRF is a useful first merge method because it can start without a training set: each result scores the sum of 1 / (k + rank) across the lists. The original paper used k = 60 and called that value "near-optimal", while also reporting that the choice was not critical (Cormack et al., 2009). Later work found RRF more sensitive to its parameter and measured a small advantage for a tuned weighted combination (Bruch et al., 2023). That alternative becomes testable after the project has a held-out set to tune against.
Chunk overlap is also an experiment, not a default to copy. Microsoft's comparison held chunking fixed at 512 tokens with 25% overlap, so it supports the retrieval comparison but does not prove that those chunk settings are best. Chroma's chunking study found up to a nine-point recall difference between strategies. In its results, overlap helped some smaller-context embedding settings but reduced token efficiency and could duplicate retrieved text. The project therefore needs a small grid over size and overlap, measured on its own transcripts, rather than one borrowed setting.
Limitations of this plan
A hand-written question set over the project's current video library is still narrow, and results on it will not transfer automatically to a very different library. YouTube's automatic captions contain recognition errors and no speaker labels. Earlier TREC Spoken Document Retrieval work found that effective retrieval was possible on automatic transcripts; that evidence predates and is separate from the TREC Podcasts track used above for entry-point judging. In the podcast collection, Clifton et al. measured 18.1% word error rate and 81.8% named-entity accuracy on a 1,600-episode sample, making names a useful evaluation slice and a likely place for keyword retrieval to help. Finally, one person writing all the questions brings one person's idea of what users ask, which is the strongest argument for adding real query logs once the tool has users.
References
- Thakur et al., BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models (2021); Muennighoff et al., MTEB (2022)
- Karpukhin et al., Dense passage retrieval for open-domain question answering (2020); Lee et al., Latent retrieval for weakly supervised open domain question answering (2019)
- Es et al., RAGAS (2023); Saad-Falcon et al., ARES (2023)
- Zheng et al., Judging LLM-as-a-judge with MT-Bench and Chatbot Arena (2023); Thakur et al., Judging the judges (2024)
- Liu et al., Lost in the middle (2023)
- Cormack, Clarke, and Büttcher, Reciprocal rank fusion (2009); Bruch, Gai, and Ingber, An analysis of fusion functions for hybrid retrieval (2023)
- Anthropic, Introducing contextual retrieval (2024); Microsoft, Azure AI Search: outperforming vector search with hybrid retrieval and reranking (2023); Chroma, Evaluating chunking strategies for retrieval (2024)
- Jones et al., TREC 2020 Podcasts track overview (2021); Garofolo et al., The TREC Spoken Document Retrieval Track: A Success Story (2000); Clifton et al., 100,000 Podcasts: A Spoken English Document Corpus (2020); Eskevich, Magdy, and Jones, New metrics for meaningful evaluation of informally structured speech retrieval (2012); Voorhees and Buckley, The effect of topic set size on retrieval experiment error (2002)