Suppose you remember a speaker comparing HNSW with IVFFlat, but you cannot remember the video title or the speaker's exact words. Keyword search works best when the question and transcript share terms; a paraphrase can be harder to find. This experiment searches the available transcripts by meaning and answers with quotes whose timestamps can seek the player to the relevant moment.
The example question on this page is illustrative. The architecture, defaults, commands, and test counts were checked against the public repository on 1 October 2026. This is a working learning project with known gaps. The page does not claim measured search quality.
How to read the diagrams
- Stored and retrieved data
- Model work: embeddings and agent decisions
- Evidence shown to the user
- Known gaps and misleading states
The concepts behind one search
A YouTube caption is a short piece of timed text. All captions together form the transcript. The ingester joins captions into chunks, passages that are long enough to carry context and small enough to search. Adjacent chunks repeat some text; that overlap keeps a sentence from disappearing at a boundary.
An embedding turns text into a list of numbers learned from patterns in language. A question and a passage about the same idea should end up close together even when their wording differs. PostgreSQL stores those lists through pgvector. Its HNSW index is a graph that follows promising neighbours instead of comparing the question with every stored vector. This makes search faster, with the trade-off that approximate search can miss a neighbour an exhaustive scan would find.
An agent is a model that can choose named tools, read their results, and choose another step. Here, each tool is ordinary Python around an embedding call or a database query. A citation is the final quote plus its start time in seconds, which the interface can turn into a seek action.
The map below separates ingestion, which happens once per video, from search, which happens for every question.
Ingest once
Timed captions1
text · start · duration
Overlapping chunks2
text · start · end
pgvector + HNSW3
1,536 numbers per chunk
Search per question
Question4
plain language → vector
Agent tools5
broad search · clip search · quotes
Return inspectable evidence
Answer + evidence6
quote · timestamp · seek
Ingesting a video
The public implementation uses four stages. Look at which information is preserved at each hand-off: the text, its timing, and its vector must stay aligned.
A YouTube video becomes searchable
Fetch timed captions
youtube-transcript-api returns caption segments. Each segment contains text, a start time and a duration.
{ text, start, duration }
Build overlapping chunks
The chunker accumulates captions up to 1,500 characters, retains up to 200 characters from the previous chunk, and drops a final fragment shorter than 100 characters.
defaults in the current public chunker
Embed in batches
text-embedding-3-small converts chunk text into 1,536-number vectors. The current embedder sends at most 512 chunks in one call.
text and vector positions remain 1:1
Replace and index
Re-ingestion updates the clip, removes its old chunks, inserts the new ones, and stores the vectors under an HNSW cosine-distance index.
one clip → many ordered chunks
The database index in the current public source is:
CREATE INDEX IF NOT EXISTS idx_transcript_chunks_embedding_hnsw
ON transcript_chunks USING hnsw (embedding vector_cosine_ops);vector_cosine_ops makes pgvector order chunks by cosine distance. Smaller distance means the directions of the two vectors are more alike. The number is a ranking signal; it is not a probability that the chunk is correct.
Where overlap breaks the timestamp
The overlap preserves words, but the current code assigns the next new caption's start time to the whole new chunk. This illustrative two-chunk trace makes the mismatch visible.
- stored chunk
- chunk 12
- text at its beginning
- “…HNSW uses more memory.”
- when those words were spoken
- 132.0 s
- stored start_time
- 120.0 s
- stored chunk
- chunk 13
- text at its beginning
- “HNSW uses more memory. IVFFlat…”
- when those words were spoken
- overlap starts at 132.0 s
- stored start_time
- 138.0 s
At 138 seconds, a new caption pushed the candidate text beyond 1,500 characters. The code finalised chunk 12, copied the last 200 characters into chunk 13, appended the new caption, and set current_start = 138.0. The copied words retain no mapping to their earlier caption, so a citation using chunk 13 can seek six seconds late in this example. A fix needs overlap at caption boundaries, or enough character-to-caption provenance to recover the first copied word's time.
Answering one question
The question “What trade-off did the speaker give for HNSW?” becomes an embedding and enters a broad nearest-neighbour search. The table below is an illustrative trace that matches the current chunk fields. Distances are included to explain ordering; the current repository query orders by cosine distance but returns chunk objects without selecting that distance into the tool payload.
- rank
- 1
- clip
- vector-index-talk
- start
- 138.0 s
- chunk text
- HNSW uses more memory but keeps query latency low…
- cosine distance
- 0.12
- rank
- 2
- clip
- database-indexing
- start
- 418.0 s
- chunk text
- The graph keeps several links for each vector…
- cosine distance
- 0.18
- rank
- 3
- clip
- vector-index-talk
- start
- 205.0 s
- chunk text
- IVFFlat partitions the vector space and probes selected lists…
- cosine distance
- 0.22
The initial tool groups those rows by clip and sends only each clip ID and its chunk text to the model. Start times and distances are absent from that first result. The agent can then search inside a promising clip and finally request cached quotes, where start times reappear.
The agent's actual three-tool loop
- User → Agent
Ask about the HNSW trade-off
- Agent → Tools
get_initial_candidates(question)
required first call
- Tools → pgvector
top 20 across all clips
ascending cosine distance
- pgvector → Tools
ranked chunk objects
- Agent → Tools
examine_clip_deeper(clip_id, question)
- Tools → pgvector
top 10 inside one clip
- Agent → Tools
get_clip_quotes(clip_id)
read the process-wide clip cache
- Tools → Agent
chronological quotes + seconds
- Agent → User
structured answer
answer · quotes · follow-ups
The trace below uses made-up content but preserves the public request and response shapes. It shows the mechanics without presenting the example as a real search result.
{
"call": {
"name": "examine_clip_deeper",
"args": {
"clip_id": "vector-index-talk",
"question": "What trade-off did the speaker give for HNSW?"
}
},
"tool_result": {
"clip_id": "vector-index-talk",
"title": "Vector index talk",
"description": "",
"transcript": "HNSW uses more memory but keeps query latency low..."
}
}After the quote tool returns text and timestamps, the model must produce the SearchResponse shape below. The instruction says the answer must use only tool information and stay under 300 characters.
{
"results": [
{
"clip_id": "vector-index-talk",
"question": "What trade-off did the speaker give for HNSW?",
"answer": "The speaker trades higher memory use for low query latency.",
"relevant_quotes": [
{
"quote": "HNSW uses more memory but keeps query latency low.",
"quote_description": "The stated index trade-off",
"quote_timestamp": 138.0
}
],
"related_questions": []
}
]
}The timestamp makes the answer inspectable. It is the retrieved chunk's start time, not the exact time of the quoted sentence inside that chunk. It does not prove that the quote entails the answer or that a better chunk was missed. Those are evaluation problems covered by the linked research note.
Why use an agent here
A fixed retrieval pipeline could embed the question, take the top chunks, and generate an answer in one route. This agent can broaden the search, inspect selected clips, and fetch quotes after it has chosen the evidence. That flexibility costs model round trips and creates more states to test.
The eight-turn cap bounds one part of that cost and prevents an endless tool loop. It does not guarantee a useful answer at turn eight. In the current graph, hitting the cap routes the latest message to the response parser; if that message contains a tool call rather than final JSON, the parser falls back to a raw-text response with no citations. A production version would need a separate, explicit “budget exhausted” result or one forced final-answer turn.
There is another boundary worth noticing. The initial tool discards start times
and distances when it groups chunks, while the deeper tool concatenates its
selected moments into one string. The quote tool restores timestamps from
_clip_cache, which is shared by the Python process. Concurrent requests can
clear or overwrite one another's cached clip entries. This keeps the model
input small, but makes ranking decisions harder to inspect and makes citation
quality depend on mutable process-wide state.
What it does not do yet
The gaps below describe the public implementation as checked on 1 October 2026. They are questions to measure and repairs to verify, rather than promised improvements.
- current gap
- Dense retrieval only
- what the user can experience
- An exact name or rare product term can rank poorly even when it appears in a transcript
- experiment or repair
- Add Postgres full-text search; fuse its ranks with vector ranks
- current gap
- No reranker
- what the user can experience
- The model receives candidates in vector-distance order, without a second relevance model
- experiment or repair
- Rerank a larger candidate set before the agent sees it
- current gap
- No evaluation set
- what the user can experience
- A search change can feel better while performing worse on held-out questions
- experiment or repair
- Label questions and compare recall, MRR, nDCG and start-time error
- current gap
- Overlap has the new caption's time
- what the user can experience
- The player can seek after the first quoted words
- experiment or repair
- Preserve timing for the copied overlap, then test start-time error
- current gap
- Quote time is chunk-granular
- what the user can experience
- A long chunk can seek before the quoted sentence even when overlap timing is fixed
- experiment or repair
- Keep caption-level provenance and return the start of the selected quote
- current gap
- Quote cache is process-wide
- what the user can experience
- Concurrent searches can replace the cache entry a later tool call expects
- experiment or repair
- Carry candidates in request or graph state, keyed by request, instead of a module global
- current gap
- Captions only
- what the user can experience
- Automatic captions can misrecognise names and provide no speaker identity
- experiment or repair
- Create caption-quality slices before choosing a speech-to-text path
The research note Evaluating Retrieval for Transcript Search turns those gaps into a test plan. No improvement is claimed until that plan produces results.
Running the public project
The current full-stack path uses Docker Compose. After cloning the repository, copy the example configuration, provide an OpenAI API key, and run:
cp .env.example .env
make upThe Streamlit interface then listens on port 8501. The public Makefile also provides make test, make test-unit, and make test-integration. The repository currently documents 36 tests: 32 unit tests across the chunker, transcriber, tools, graph and schemas, plus four HTTP integration tests. Those counts were checked from the public test listing; this portfolio review did not execute the separate repository.
Sources
- Public repository: README and system overview, checked 1 October 2026
- Public source: chunk construction, embedding batches, and three agent tools
- Public source: agent graph and eight-turn cap, vector queries, and developer commands
- pgvector: HNSW and distance operators