Skip to content

Project

Semantic Video Search

An open-source search experiment for video transcripts: YouTube captions become overlapping chunks in pgvector, a LangGraph agent chooses retrieval tools, and each answer points back to quoted moments in the video.

experimentPublished 1 Jun 2026Updated 2 Oct 20265 min read
AI Engineering · RAG · Agents · Semantic Search · Vector Search
A question becoming an embedding, passing through ranked transcript chunks, and returning a timestamped video quote.
Status
experiment
Stack
Python · FastAPI · LangGraph · PostgreSQL · pgvector · Streamlit · OpenAI · Anthropic
Repository
GitHub
On this page (8)

Suppose you remember a speaker comparing HNSW with IVFFlat, but you cannot remember the video title or the speaker's exact words. Keyword search works best when the question and transcript share terms; a paraphrase can be harder to find. This experiment searches the available transcripts by meaning and answers with quotes whose timestamps can seek the player to the relevant moment.

The example question on this page is illustrative. The architecture, defaults, commands, and test counts were checked against the public repository on 1 October 2026. This is a working learning project with known gaps. The page does not claim measured search quality.

How to read the diagrams

  • Stored and retrieved data
  • Model work: embeddings and agent decisions
  • Evidence shown to the user
  • Known gaps and misleading states

A YouTube caption is a short piece of timed text. All captions together form the transcript. The ingester joins captions into chunks, passages that are long enough to carry context and small enough to search. Adjacent chunks repeat some text; that overlap keeps a sentence from disappearing at a boundary.

An embedding turns text into a list of numbers learned from patterns in language. A question and a passage about the same idea should end up close together even when their wording differs. PostgreSQL stores those lists through pgvector. Its HNSW index is a graph that follows promising neighbours instead of comparing the question with every stored vector. This makes search faster, with the trade-off that approximate search can miss a neighbour an exhaustive scan would find.

An agent is a model that can choose named tools, read their results, and choose another step. Here, each tool is ordinary Python around an embedding call or a database query. A citation is the final quote plus its start time in seconds, which the interface can turn into a seek action.

The map below separates ingestion, which happens once per video, from search, which happens for every question.

Ingest once

Timed captions1

text · start · duration

Overlapping chunks2

text · start · end

pgvector + HNSW3

1,536 numbers per chunk

embed with the same model

Search per question

Question4

plain language → vector

Agent tools5

broad search · clip search · quotes

Return inspectable evidence

Answer + evidence6

quote · timestamp · seek

The stored chunk is the bridge: ingestion gives it text, time and an embedding; search ranks it and the answer preserves its quote and time.

Ingesting a video

The public implementation uses four stages. Look at which information is preserved at each hand-off: the text, its timing, and its vector must stay aligned.

A YouTube video becomes searchable

  1. Fetch timed captions

    youtube-transcript-api returns caption segments. Each segment contains text, a start time and a duration.

    { text, start, duration }

  2. Build overlapping chunks

    The chunker accumulates captions up to 1,500 characters, retains up to 200 characters from the previous chunk, and drops a final fragment shorter than 100 characters.

    defaults in the current public chunker

  3. Embed in batches

    text-embedding-3-small converts chunk text into 1,536-number vectors. The current embedder sends at most 512 chunks in one call.

    text and vector positions remain 1:1

  4. Replace and index

    Re-ingestion updates the clip, removes its old chunks, inserts the new ones, and stores the vectors under an HNSW cosine-distance index.

    one clip → many ordered chunks

The database index in the current public source is:

backend/app/db/session.py (public excerpt)
CREATE INDEX IF NOT EXISTS idx_transcript_chunks_embedding_hnsw
ON transcript_chunks USING hnsw (embedding vector_cosine_ops);

vector_cosine_ops makes pgvector order chunks by cosine distance. Smaller distance means the directions of the two vectors are more alike. The number is a ranking signal; it is not a probability that the chunk is correct.

Where overlap breaks the timestamp

The overlap preserves words, but the current code assigns the next new caption's start time to the whole new chunk. This illustrative two-chunk trace makes the mismatch visible.

illustrative chunk boundary
stored chunk
chunk 12
text at its beginning
“…HNSW uses more memory.”
when those words were spoken
132.0 s
stored start_time
120.0 s
stored chunk
chunk 13
text at its beginning
“HNSW uses more memory. IVFFlat…”
when those words were spoken
overlap starts at 132.0 s
stored start_time
138.0 s

At 138 seconds, a new caption pushed the candidate text beyond 1,500 characters. The code finalised chunk 12, copied the last 200 characters into chunk 13, appended the new caption, and set current_start = 138.0. The copied words retain no mapping to their earlier caption, so a citation using chunk 13 can seek six seconds late in this example. A fix needs overlap at caption boundaries, or enough character-to-caption provenance to recover the first copied word's time.

Answering one question

The question “What trade-off did the speaker give for HNSW?” becomes an embedding and enters a broad nearest-neighbour search. The table below is an illustrative trace that matches the current chunk fields. Distances are included to explain ordering; the current repository query orders by cosine distance but returns chunk objects without selecting that distance into the tool payload.

illustrative top results
rank
1
clip
vector-index-talk
start
138.0 s
chunk text
HNSW uses more memory but keeps query latency low…
cosine distance
0.12
rank
2
clip
database-indexing
start
418.0 s
chunk text
The graph keeps several links for each vector…
cosine distance
0.18
rank
3
clip
vector-index-talk
start
205.0 s
chunk text
IVFFlat partitions the vector space and probes selected lists…
cosine distance
0.22

The initial tool groups those rows by clip and sends only each clip ID and its chunk text to the model. Start times and distances are absent from that first result. The agent can then search inside a promising clip and finally request cached quotes, where start times reappear.

The agent's actual three-tool loop

  1. User → Agent

    Ask about the HNSW trade-off

  2. Agent → Tools

    get_initial_candidates(question)

    required first call

  3. Tools → pgvector

    top 20 across all clips

    ascending cosine distance

  4. pgvector → Tools

    ranked chunk objects

  5. Agent → Tools

    examine_clip_deeper(clip_id, question)

  6. Tools → pgvector

    top 10 inside one clip

  7. Agent → Tools

    get_clip_quotes(clip_id)

    read the process-wide clip cache

  8. Tools → Agent

    chronological quotes + seconds

  9. Agent → User

    structured answer

    answer · quotes · follow-ups

The model chooses the calls, while Python owns the search and quote retrieval. The graph stops when the model emits no tool call or reaches eight model turns.

The trace below uses made-up content but preserves the public request and response shapes. It shows the mechanics without presenting the example as a real search result.

Sketch: one agent search trace
{
  "call": {
    "name": "examine_clip_deeper",
    "args": {
      "clip_id": "vector-index-talk",
      "question": "What trade-off did the speaker give for HNSW?"
    }
  },
  "tool_result": {
    "clip_id": "vector-index-talk",
    "title": "Vector index talk",
    "description": "",
    "transcript": "HNSW uses more memory but keeps query latency low..."
  }
}

After the quote tool returns text and timestamps, the model must produce the SearchResponse shape below. The instruction says the answer must use only tool information and stay under 300 characters.

Sketch: final SearchResponse
{
  "results": [
    {
      "clip_id": "vector-index-talk",
      "question": "What trade-off did the speaker give for HNSW?",
      "answer": "The speaker trades higher memory use for low query latency.",
      "relevant_quotes": [
        {
          "quote": "HNSW uses more memory but keeps query latency low.",
          "quote_description": "The stated index trade-off",
          "quote_timestamp": 138.0
        }
      ],
      "related_questions": []
    }
  ]
}

The timestamp makes the answer inspectable. It is the retrieved chunk's start time, not the exact time of the quoted sentence inside that chunk. It does not prove that the quote entails the answer or that a better chunk was missed. Those are evaluation problems covered by the linked research note.

Why use an agent here

A fixed retrieval pipeline could embed the question, take the top chunks, and generate an answer in one route. This agent can broaden the search, inspect selected clips, and fetch quotes after it has chosen the evidence. That flexibility costs model round trips and creates more states to test.

The eight-turn cap bounds one part of that cost and prevents an endless tool loop. It does not guarantee a useful answer at turn eight. In the current graph, hitting the cap routes the latest message to the response parser; if that message contains a tool call rather than final JSON, the parser falls back to a raw-text response with no citations. A production version would need a separate, explicit “budget exhausted” result or one forced final-answer turn.

There is another boundary worth noticing. The initial tool discards start times and distances when it groups chunks, while the deeper tool concatenates its selected moments into one string. The quote tool restores timestamps from _clip_cache, which is shared by the Python process. Concurrent requests can clear or overwrite one another's cached clip entries. This keeps the model input small, but makes ranking decisions harder to inspect and makes citation quality depend on mutable process-wide state.

What it does not do yet

The gaps below describe the public implementation as checked on 1 October 2026. They are questions to measure and repairs to verify, rather than promised improvements.

known gaps
current gap
Dense retrieval only
what the user can experience
An exact name or rare product term can rank poorly even when it appears in a transcript
experiment or repair
Add Postgres full-text search; fuse its ranks with vector ranks
current gap
No reranker
what the user can experience
The model receives candidates in vector-distance order, without a second relevance model
experiment or repair
Rerank a larger candidate set before the agent sees it
current gap
No evaluation set
what the user can experience
A search change can feel better while performing worse on held-out questions
experiment or repair
Label questions and compare recall, MRR, nDCG and start-time error
current gap
Overlap has the new caption's time
what the user can experience
The player can seek after the first quoted words
experiment or repair
Preserve timing for the copied overlap, then test start-time error
current gap
Quote time is chunk-granular
what the user can experience
A long chunk can seek before the quoted sentence even when overlap timing is fixed
experiment or repair
Keep caption-level provenance and return the start of the selected quote
current gap
Quote cache is process-wide
what the user can experience
Concurrent searches can replace the cache entry a later tool call expects
experiment or repair
Carry candidates in request or graph state, keyed by request, instead of a module global
current gap
Captions only
what the user can experience
Automatic captions can misrecognise names and provide no speaker identity
experiment or repair
Create caption-quality slices before choosing a speech-to-text path

The research note Evaluating Retrieval for Transcript Search turns those gaps into a test plan. No improvement is claimed until that plan produces results.

Running the public project

The current full-stack path uses Docker Compose. After cloning the repository, copy the example configuration, provide an OpenAI API key, and run:

Public quick-start commands
cp .env.example .env
make up

The Streamlit interface then listens on port 8501. The public Makefile also provides make test, make test-unit, and make test-integration. The repository currently documents 36 tests: 32 unit tests across the chunker, transcriber, tools, graph and schemas, plus four HTTP integration tests. Those counts were checked from the public test listing; this portfolio review did not execute the separate repository.

Sources