Skip to content

Research Entry

Evaluating Retrieval for Transcript Search

A measurement plan for transcript search: separate retrieval from generation, build a fair test set, calibrate automated judges, and compare dense, hybrid, reranked, and timestamp-aware search.

researchingPublished 15 Jul 2026Updated 2 Oct 20269 min read
RAG · Evaluation · Information Retrieval · AI Engineering
An explicitly illustrative ranked list of transcript chunks flowing into recall, MRR, and nDCG calculations.
Status
researching
Question
How should retrieval quality be measured for timestamped transcript search, and which changes are most likely to improve it?
Last updated
2 Oct 2026
On this page (8)

My semantic video search project finds short passages from video transcripts and gives them to a model to answer a question. It currently ranks those passages by one semantic search method and has no labelled test set. That makes every proposed improvement, from keyword search to a reranker or a different chunk size, an untested idea.

This note turns those ideas into a measurement plan: what to measure, how to build a fair test set, how far to trust an automated judge, and which experiments published evidence suggests running first.

How to read the diagrams

  • Retrieval: finding the right chunks
  • Generation: what the model does with them
  • People: labels, judgements, and the questions themselves
  • Published evidence that a change helps
  • Traps that make an evaluation lie

The system being measured

Retrieval-augmented generation (RAG) first finds source material, then gives it to a model that writes an answer. A transcript is divided into chunks, short passages small enough to retrieve and quote. An embedding represents a chunk or question as numbers so nearby meanings can be found by vector search.

The retriever returns a ranked list of chunks. The generator reads those chunks and writes the answer. BM25 is a keyword-ranking method that rewards query terms while reducing the influence of very common words. Hybrid search combines keyword and vector results. Reciprocal rank fusion (RRF) is one way to do that: it gives a result points from its position in each ranked list, then adds the points. A reranker reads the question and each candidate together, then produces a more expensive relevance score for a smaller candidate set.

Fixed evaluation input

Questions1

human-written + reference timestamps

run each retriever on the same questions

System under test

Retriever variant2

dense · hybrid · reranked

Ranked chunks3

top 5 · 10 · 20

Judgement

Relevance labels4

0 · 1 · 2 · 3 + correct start

Measurement and decision

Metrics5

recall · MRR · nDCG · time error

Decision6

keep · reject · investigate a slice

Labels turn ranked chunks into measurements. A change is kept only when it improves the chosen metrics on held-out questions without breaking latency or timestamp accuracy.

Relevance labels say how well a chunk answers one question. This plan uses 0 for irrelevant, 1 for related, 2 for useful but incomplete, and 3 for the exact evidence. An LLM judge is a model asked to assign those labels; human labels remain the reference used to test whether the automated judge is dependable.

Retrieval and generation are different measurements

A RAG answer can be wrong for two unrelated reasons: the right passage was never retrieved, or it was retrieved and the model misused it. Measuring only the final answer mixes the two, so a change that improves retrieval can look like no change at all if the model then ignores the extra context. In Lost in the Middle, performance was often strongest when the needed information appeared near the beginning or end of the context and weaker when it appeared in the middle. The size of that effect depended on the model and task, but the lesson is useful here: better recall does not automatically produce better answers.

So the plan measures the retriever on its own first, with no model involved, and the generator second.

What to measure

retrieval metrics
metric
Recall@k
what it measures
Share of the relevant chunks that appear in the top k
when it's the right one
When a model will read all k chunks anyway: did the evidence make it in?
metric
Hit rate@k
what it measures
Share of questions with at least one relevant chunk in the top k (DPR's 'top-20 accuracy')
when it's the right one
When one good chunk is enough to answer
metric
MRR
what it measures
Mean of 1 / rank of the first relevant result (0 if none)
when it's the right one
When the first good result is what matters, as with the agent's first search
metric
nDCG@k
what it measures
Rewards relevant results near the top, and handles graded relevance (perfect vs partly relevant)
when it's the right one
The standard summary: BEIR and MTEB both report nDCG@10 as the main number
metric
Start-time error
what it measures
Distance in seconds between the retrieved chunk's start and the moment the answer begins
when it's the right one
Specific to timestamped media: a right chunk that seeks to the wrong moment is half a failure

Here is one worked example. These five chunks and scores are illustrative; they are not results from the project. Suppose the evaluation pool contains exactly three chunks labelled relevant to the question, and all three appear below. A relevance grade of 1, 2 or 3 counts as relevant for recall and MRR. That broad threshold answers a practical retrieval question: did the candidate set contain material that could help the next stage? The stricter grade 2-or-higher result will also be reported as a sensitivity check, because a merely related chunk may not contain enough evidence to answer.

question · What trade-off did the speaker give for HNSW?
rank
1
retrieved chunk
How IVFFlat partitions the vector space
grade
0 · irrelevant
first relevant?
no
rank
2
retrieved chunk
HNSW uses more memory but gives strong recall and low query latency
grade
3 · exact answer
first relevant?
yes
rank
3
retrieved chunk
The graph stores several links for each vector
grade
1 · related
first relevant?
already found
rank
4
retrieved chunk
How cosine distance is computed
grade
0 · irrelevant
first relevant?
not applicable
rank
5
retrieved chunk
Index builds can take longer as the collection grows
grade
2 · useful
first relevant?
already found

Recall@5 asks how many known relevant chunks appeared: 3 retrieved ÷ 3 relevant = 1.00. It ignores order, so putting the exact answer at rank 2 rather than rank 5 makes no difference.

MRR means mean reciprocal rank. For this single question, the first relevant result is rank 2, so its reciprocal rank is 1 ÷ 2 = 0.50. Across a test set, average that value for every question; a question with no relevant result contributes zero.

nDCG@5 keeps the grades and the order. This example uses gain 2^grade − 1 and discount log₂(rank + 1). Its discounted cumulative gain is:

Sketch: illustrative DCG@5 calculation
rank 2: (2³ − 1) / log₂(3) = 4.417
rank 3: (2¹ − 1) / log₂(4) = 0.500
rank 5: (2² − 1) / log₂(6) = 1.161
DCG@5 = 6.077

The ideal order is grades 3, 2, 1 at ranks 1, 2, 3. Its IDCG@5 is 9.393, so nDCG@5 = 6.077 ÷ 9.393 = 0.647. Normalising by the ideal makes questions with different numbers of relevant chunks comparable on a 0–1 scale.

The last row matters for transcripts. The TREC Podcasts track, which evaluated search over 100,000 podcast episodes, graded segments partly on whether they are "an ideal entry point for a human listener", and earlier speech-retrieval work (CLEF's mGAP measure) penalised the distance between the retrieved start and the true start, up to a limit of 150 seconds (Eskevich et al., 2012). The project's timestamps are the whole point of its answers, so start-time error gets measured alongside ranking quality.

For generation, the useful measurements are the ones RAGAS popularised: faithfulness (the share of the answer's claims that the retrieved context supports), answer relevance, and context relevance. In the RAGAS paper's own validation, the automated scores agreed with human annotators 95% of the time for faithfulness but only 78% for answer relevance and 70% for context relevance, which they called "the hardest quality dimension to evaluate." Those are results for that paper's judge and datasets, not a reliability guarantee. Faithfulness is the first automated signal to test against this project's human labels.

Building a test set that doesn't flatter one retriever

The test set decides the winner more often than the retriever does. Two traps are well documented.

Questions written while looking at the passage favour keyword search. The ORQA authors pointed out that in SQuAD, question writers "are also prompted with a specific piece of evidence for the answer, leading to artificially large lexical overlap between the question and evidence." The DPR paper measured the effect: on SQuAD, BM25 retrieved an answer-bearing passage in the top 20 for 68.8% of questions against DPR's 63.2%, while on Natural Questions, where real users wrote the questions without seeing the answer, the order reversed, 59.1% against 78.4% (Karpukhin et al., 2020). A test set generated by asking a model to write questions about a chunk has the same shape, so it will overstate how well keyword matching works.

Unjudged is not irrelevant. Most test sets label a few passages per question and treat everything else as wrong. BEIR found that 31.8% of TAS-B's top-10 results were unjudged, against 6.4% for BM25, because the original judgement pools came from keyword-based systems. In a separate TREC-COVID reannotation experiment, ANCE's nDCG@10 rose from 0.654 to 0.735 after missing judgements were filled in. Those numbers come from different retrievers and experiments; together they show why each compared system must contribute candidates to the judgement pool.

The plan for this project's test set follows from both:

  • Begin with roughly 50 questions, written by people who watched the videos rather than by a model reading chunks. Treat that as a pilot, not a magic minimum. TREC evaluations have often used about 50 topics, while Voorhees and Buckley show that topic-set size changes experiment error. Add questions, or repeat the comparison over resampled question sets, until the uncertainty is small enough for the decision being made. ARES's separate suggestion of about 150 labelled datapoints applies to judge calibration, not to the number of unique questions (Voorhees and Buckley, 2002; Saad-Falcon et al., 2023).
  • Graded relevance on four levels, as TREC Deep Learning uses: 3 contains the exact answer, 2 contains it less clearly, 1 is related but doesn't answer, 0 is irrelevant.
  • For every question, the correct start time in seconds.
  • Pooled judging: every system being compared contributes its top results to the pool that gets labelled, so no retriever is scored against labels that only another retriever could have found.

How far to trust an LLM judge

Labelling every result by hand doesn't scale, so an LLM judge will do most of the labelling, and it has known biases. A widely cited study, Judging LLM-as-a-Judge, found that GPT-4 agreed with human preferences 85% of the time on clear-cut comparisons, slightly more than humans agreed with each other (81%). The same paper found that the order of the two answers changed the verdict: GPT-4 gave the same judgement when the answers were swapped only 65% of the time.

judge biases measured in Zheng et al. (2023)
bias
Position
measurement
Consistent verdict after swapping answer order: GPT-4 65%, GPT-3.5 46.2%, Claude-v1 23.8%
mitigation used in this plan
Judge each pair twice in both orders; count a win only if both agree
bias
Verbosity
measurement
A padded, repetitive answer fooled GPT-3.5 and Claude-v1 in 91.3% of cases, GPT-4 in 8.7%
mitigation used in this plan
Grade relevance of a chunk to a question, not preference between answers
bias
Self-preference
measurement
GPT-4 gave itself a 10% higher win rate, Claude-v1 25% (the authors note this can't prove bias)
mitigation used in this plan
Use a different model to judge than the one that generates
bias
No reference
measurement
On math questions, judge failures dropped from 14 of 20 to 3 of 20 when given a reference answer
mitigation used in this plan
Give the judge the human-written reference answer and timestamp

Agreement percentage alone can hide a problem: later work found that judges can agree often while still assigning materially different scores, and that some judges lean toward generous grades (Thakur et al., 2024). Two people will independently label a calibration subset and adjudicate disagreements before it becomes the reference. The first calibration pool will contain at least 150 query-and-chunk judgements, following ARES's input guidance, and will grow if rare grades or question types have too few examples. The model judge will then be compared with that reference using weighted Cohen's kappa, because the four labels are ordered and confusing 3 with 2 is less severe than confusing 3 with 0. Raw agreement, the confusion matrix, and per-grade precision and recall will be shown beside it; one summary coefficient cannot reveal which errors the judge makes.

What to try first, and why

Each experiment below changes one thing, and each is here because published work measured a real effect on a comparable problem.

Anthropic's contextual retrieval: share of queries where the right chunk missed the top 20

Retrieval failure rate

Adding a short model-written context to each chunk, then contextual BM25, then a reranker cut failures by 35%, 49%, and 67%.

y: percent of queries (1 − recall@20)

Anthropic, “Contextual Retrieval” (September 2024), across codebases, fiction, and papers.

Microsoft's hybrid retrieval measurements across four customer datasets

Ranking quality

On queries that were mostly exact keywords, vector search scored 11.7 against keyword search's 79.2.

y: NDCG@3

Azure AI Search, “Outperforming vector search with hybrid retrieval and reranking” (September 2023).
planned experiments
project gap → experiment
No evaluation set → build and freeze the first labelled split
why this experiment
Every later comparison needs the same questions, labels and metric code
decision metrics
judge agreement first; then baseline recall@20, MRR, nDCG@10 and start-time error
project gap → experiment
Dense retrieval only → add full-text search and merge with RRF
why this experiment
BEIR found BM25 remained a strong out-of-domain baseline; Microsoft's four-dataset study scored hybrid above vector alone (48.4 vs 43.8 nDCG@3)
decision metrics
nDCG@10 and recall@20; slice names, jargon and exact quotes
project gap → experiment
No reranker → cross-encode the top 50 candidates
why this experiment
BEIR reported BM25 plus a cross-encoder above BM25 on 16 of 18 datasets, with an 11% average gain and about 450 ms GPU latency in that setup
decision metrics
nDCG@10, MRR and added p50/p95 latency
project gap → experiment
Dense retrieval only → add video and section context to each chunk
why this experiment
Anthropic reported a 35% reduction in top-20 retrieval failures from contextual embeddings on its evaluated corpora
decision metrics
recall@20; compare separately from hybrid so the cause is identifiable
project gap → experiment
Chunk start times → fix overlap timestamps, then vary size and overlap
why this experiment
Timestamp correctness is a known project bug. Microsoft used 512-token chunks with 25% overlap for all four customer datasets; Chroma found that the best size and overlap depended on the embedding model and that overlap can add redundant tokens
decision metrics
start-time error first; then recall@20 and retrieved-token efficiency for each chunking variant
project gap → experiment
Captions only → create an ASR and named-entity slice
why this experiment
Automatic captions can fail on names, which is also where keyword retrieval may help
decision metrics
all retrieval metrics split by caption source and named-entity presence

RRF is a useful first merge method because it can start without a training set: each result scores the sum of 1 / (k + rank) across the lists. The original paper used k = 60 and called that value "near-optimal", while also reporting that the choice was not critical (Cormack et al., 2009). Later work found RRF more sensitive to its parameter and measured a small advantage for a tuned weighted combination (Bruch et al., 2023). That alternative becomes testable after the project has a held-out set to tune against.

Chunk overlap is also an experiment, not a default to copy. Microsoft's comparison held chunking fixed at 512 tokens with 25% overlap, so it supports the retrieval comparison but does not prove that those chunk settings are best. Chroma's chunking study found up to a nine-point recall difference between strategies. In its results, overlap helped some smaller-context embedding settings but reduced token efficiency and could duplicate retrieved text. The project therefore needs a small grid over size and overlap, measured on its own transcripts, rather than one borrowed setting.

Limitations of this plan

A hand-written question set over the project's current video library is still narrow, and results on it will not transfer automatically to a very different library. YouTube's automatic captions contain recognition errors and no speaker labels. Earlier TREC Spoken Document Retrieval work found that effective retrieval was possible on automatic transcripts; that evidence predates and is separate from the TREC Podcasts track used above for entry-point judging. In the podcast collection, Clifton et al. measured 18.1% word error rate and 81.8% named-entity accuracy on a 1,600-episode sample, making names a useful evaluation slice and a likely place for keyword retrieval to help. Finally, one person writing all the questions brings one person's idea of what users ask, which is the strongest argument for adding real query logs once the tool has users.

References