A semantic cache answers a new request with an answer created for an earlier, different request. It decides that the two requests are close enough in meaning. A correct match avoids the model call. A false match returns a polished answer to the wrong question, which can be harder to notice than an error.
Use the provider's exact-prefix prompt cache where its contract fits; it reuses prompt computation while the model still produces a new answer. Use a semantic response cache only for narrow, repetitive, short tasks, inside one tenant boundary, after measuring false hits. Do not apply it to an agent's coordinator turns or to a whole multi-turn conversation.
How to read the diagrams
- The request and the cache's records
- A correct reuse: the stored answer fits
- A false hit: the stored answer is for a different question
- A provider call: the miss path, where tokens cost money
- Tenant and policy boundaries
Three kinds of reuse
Exact response cache
- Matches
- A byte-identical request inside the same complete cache key
- Saves
- The whole call
- Risk
- Low only when tenant, instructions, model policy, permissions, and freshness are also part of eligibility
- Example
- Cloudflare AI Gateway caches identical requests only
Provider prompt cache
- Matches
- An identical prefix of the prompt
- Saves
- Input processing: in September 2026, GPT-5.6 Terra cached input is $0.20/M versus $2.00/M uncached; Anthropic rates vary by model and cache duration
- Risk
- It does not reuse the prior final answer; provider eligibility and expiry rules still apply
- Example
- Supported OpenAI and Anthropic models
Semantic cache
- Matches
- A different request with a close embedding
- Saves
- The whole call, output included
- Risk
- Serving the answer to a different question
- Example
- GPTCache, Redis LangCache, gateway plugins
Exact and provider-prefix caches reuse work for identical bytes. Identical input is not enough by itself: an exact response can still be stale or belong to another tenant or prompt version, so it needs the same policy boundaries. The semantic cache crosses an additional boundary and decides that different text deserves the same answer. The rest of this chapter is about making that decision visible and measurable.
How text becomes a similarity score
An embedding turns text into a vector, a list of numbers that places the text on a model's learned map. Nearby directions on that map often express related ideas. Cosine similarity compares the angle between two vectors: 1 means the directions are identical, lower values mean they point farther apart. The score describes the embedding model's geometry; it does not prove that two requests have the same correct answer.
A nearest-neighbour search asks which stored vectors sit closest to the new vector. A threshold is the minimum score the application will accept as a cache hit. The map analogy stops here: names, dates, negation, and product boundaries can change the correct answer without moving the vector very far.
The following scores are illustrative. The cached request is “Can you move my demo to Thursday?” and this example uses a threshold of 0.92.
- new request
- Please reschedule my demo for Thursday
- similarity
- 0.96
- cache decision
- hit
- should reuse?
- yes
- new request
- Move my onboarding call to Thursday
- similarity
- 0.94
- cache decision
- hit
- should reuse?
- no: different appointment
- new request
- Can you move Jordan's demo to Thursday?
- similarity
- 0.92
- cache decision
- hit
- should reuse?
- no: different person
- new request
- Can you move my demo to Friday?
- similarity
- 0.91
- cache decision
- miss
- should reuse?
- no: different date
- new request
- Can you cancel my demo on Thursday?
- similarity
- 0.89
- cache decision
- miss
- should reuse?
- no: opposite action
- new request
- What will the weather be on Thursday?
- similarity
- 0.41
- cache decision
- miss
- should reuse?
- no: unrelated
What one lookup does
The running example is a use case that suggests a short reply to an inbound customer message. Here is one lookup, with the filters that decide which stored answers are even eligible:
suggest_reply · one cache lookup
Embed the request
The user's message is turned into a vector with the same embedding model used for the stored entries.
illustrative: text-embedding-3-small · 1,536 dimensions
Filter to eligible entries
Only entries inside the same tenant and policy boundary can match, and only while they are fresh.
illustrative: org_42 · suggest_reply · prompt h:9c2e… · policy v4 · not expired
Find neighbours above the floor
The threshold is inside the query. The closest entry is still a miss when its score is too low.
illustrative: cosine similarity ≥ 0.92 · top 5
Apply the serving policy
Similarity is necessary but not sufficient. The entry must also have passed quality checks when it was written.
quality ok · not sampled for regeneration today
Serve the stored answer
The request does not reach the generation model. Record the source entry and score so the decision can be audited.
illustrative: hit cache_17 · score 0.94
In SQL, with pgvector, the lookup looks like this:
SELECT id, response, 1 - (embedding <=> $1) AS similarity
FROM llm_cache
WHERE org_id = $2 -- tenant: never optional
AND use_case = $3
AND prompt_hash = $4 -- a new system prompt starts a new cache
AND embedding_model = $5
AND policy_version = $6
AND locale = $7
AND quality_status = 'approved'
AND deleted_at IS NULL
AND fresh_until > now()
AND (embedding <=> $1) <= 1 - $8 -- $8 = similarity floor
ORDER BY embedding <=> $1
LIMIT 5;The distance condition belongs in the WHERE clause. Without it, the query returns the nearest neighbour even when the nearest neighbour is not similar at all, and correctness then depends on every caller remembering to check the score.
The threshold is the product
Similarity scores do not know the cost of a wrong answer. That makes the threshold a product decision rather than a model default. Here is the running request against four stored requests:
| stored request | similarity | same answer? |
|---|---|---|
| Please reschedule my demo for Thursday | 0.96 | yes: same intent, same day |
| Could you move my demo to Friday? | 0.93 | no: a different day |
| Can you cancel my demo on Thursday? | 0.91 | no: the opposite request |
| What time is my demo on Thursday? | 0.88 | no, and below the floor |
A floor of 0.90 serves the first row, which is right, but also the second and third, which are wrong. Raising the floor to 0.935 fixes this tiny example but cannot prove the next dataset will behave the same way. Small words that change the answer, such as Friday, cancel, and not, may move an embedding less than expected.
Measure the trade-off on labelled requests from the actual use case. Precision asks: of the requests served from cache, how many were correct? Recall asks: of all requests that could safely have reused an answer, how many did the cache find? A higher threshold normally raises precision and lowers recall.
| threshold | correct hits | wrong hits | safe matches missed | precision | recall |
|---|---|---|---|---|---|
| 0.88 | 42 | 12 | 8 | 77.8% | 84% |
| 0.92 | 29 | 4 | 21 | 87.9% | 58% |
| 0.96 | 12 | 1 | 38 | 92.3% | 24% |
Published measurements show why a threshold needs local evaluation. In an appendix experiment on English WildChat samples, the InstCache authors used GPTCache with its Albert-small embedding model and used DeepSeek V3 to judge whether matched instructions would have similar answers:
GPTCache on WildChat conversations (InstCache, 2025)
Share of requests served from cache
Raising the floor lowers the hit rate, as expected.
Share of those hits that were wrong
It doesn't lower the error rate: about a third of hits were mismatches at both floors.
y: percent
TweakLLM measured GPTCache on a question-pairs dataset. At threshold 0.70, its reported precision was about 0.90; with an ALBERT re-ranker at threshold 0.97, precision reached 0.97 while recall fell to about 0.20. vCache compared its dynamic threshold with the static-threshold and other baselines in its own workloads, reporting up to 12.5× more cache hits and up to 26× fewer errors. Those ratios are comparisons inside that experiment, not production guarantees for another embedding model or dataset.
Set the floor per use case. Classification into a small fixed set may tolerate a trade-off that is unacceptable for writing a customer reply. Build should-match and should-not-match pairs, then compute precision and recall at each candidate floor. If the measured false-hit cost is still too high, bypass semantic reuse for that use case.
A product can scope semantic-cache eligibility to a tenant and use case, then evaluate candidate thresholds in shadow mode before serving reviewed matches. The example is a rollout pattern, not a statement about a particular product's plans or active controls.
The cache key is a security boundary
Every filter in the lookup above is there for correctness, and the tenant filter is also there for security. A cache that matches across tenants will eventually serve one customer's data to another, because a similar question from a different company is still a similar question.
Off-the-shelf tools are explicit about where their defaults draw the line. LiteLLM's documentation warns that end users behind one virtual key share a semantic bucket by default, which allows one user's response, including tool calls, to match another user's similar prompt. Portkey's semantic matching ignores the system prompt, so two use cases with different instructions can answer each other's requests; its semantic-cache feature is listed for Enterprise plans as of September 2026. Redis recommends "hard metadata boundaries (tenant, locale, model version, safety flags)". CacheAttack showed that an attacker who can write to a shared semantic cache can craft entries that hijack other users' responses, reporting an 86% attack hit rate in its experiments. Locality makes a semantic cache useful, but it does not provide the collision resistance of a cryptographic key.
One illustrative row makes the boundary concrete:
- field
- org_id
- example
- org_42
- if omitted from eligibility
- another customer's answer can cross the tenant boundary
- field
- use_case
- example
- suggest_reply
- if omitted from eligibility
- a classifier result can serve a writing request
- field
- prompt_hash
- example
- 9c2e…
- if omitted from eligibility
- old instructions keep answering after a prompt change
- field
- embedding_model
- example
- text-embedding-3-small
- if omitted from eligibility
- scores compare vectors from incompatible spaces
- field
- policy_version
- example
- 4
- if omitted from eligibility
- an entry approved under an older safety or quality rule remains eligible
- field
- locale
- example
- en-GB
- if omitted from eligibility
- the answer can use the wrong language or formatting convention
- field
- fresh_until
- example
- illustrative: day 10
- if omitted from eligibility
- date-sensitive answers can live forever
- field
- generation_model
- example
- recorded, policy decides whether it filters
- if omitted from eligibility
- a model upgrade may keep serving older-model answers without an explicit decision
- field
- request · embedding · response
- example
- encrypted payload + vector + answer
- if omitted from eligibility
- there is nothing to compare or return; these are stored data, not boundary fields
- field
- quality status · deleted_at
- example
- approved · NULL
- if omitted from eligibility
- rejected or deleted entries can be served
Database row-level security can reinforce tenant isolation. Tests should attempt a cross-tenant lookup and prove it returns no row. Whether the generation model is part of eligibility is an explicit product decision: including it drops all old-model hits after an upgrade; omitting it saves more calls but keeps serving older-model answers until another boundary changes.
Deduplicating concurrent identical misses (so ten simultaneous copies of the same request make one provider call) is worth doing, and it has the same rule: the deduplication key includes the tenant, so it never merges requests from two organisations.
Do not reuse agent turns as semantic-cache answers
The strongest warning in the published material concerns agents. LiteLLM's docs explain that in an agent loop "consecutive turns are nearly identical text and their embeddings are ~0.99 similar. At any practical threshold, every turn matches the previous turn's cached entry … which typically shows up as an agent repeating the same tool call over and over. Raising similarity_threshold does not reliably fix this."
The reason is structural. In a multi-turn conversation, the new information is a small addition at the end of a long shared history, and an embedding of the whole history is dominated by the shared part. The cache sees "the same conversation" and returns the previous step's answer.
Structured outputs that feed code need the same caution. A semantically close request can require a different identifier, enum member, or tool argument while still producing a high embedding score. Reusing the old JSON can pass schema validation and execute the wrong action, so these calls bypass response reuse.
Keep semantic response reuse to independently answerable requests. An agent can still use exact caches for deterministic tool results and a provider's prefix cache for stable instructions and history. Those mechanisms do not substitute a previous coordinator decision merely because the new turn embeds nearby.
The index: exact scan or approximate search
pgvector is a PostgreSQL extension that stores vectors and calculates their distance. An exact search compares the query vector with every eligible row. An approximate nearest-neighbour index searches a smaller part of the space to respond faster, accepting that it may miss a true neighbour. HNSW, short for hierarchical navigable small world, is one such graph-based index; IVFFlat is another.
When tenant and use-case filters leave a small candidate set, a B-tree index on those fields can narrow the rows before an exact distance calculation. Exact search never loses recall because of an approximate index, so start there and measure before adding HNSW.
Approximate indexes become useful when the eligible set is large, and they have a sharp edge with filters. The pgvector README explains that filtering is applied after an approximate scan. If the scan visits 40 candidates and a filter keeps 10% of the overall rows, it keeps about four candidates on average. Iterative scans can search farther when filters remove results, while partial indexes or partitioning can narrow the indexed population. Measure recall by comparing approximate results with an exact scan on the same labelled queries.
Freshness, deletion, and what gets stored
Stored answers age. A time to live, or TTL, is how long an entry stays eligible. The illustrative ten-day window in the lookup is an example, not a production setting. A scheduled retention job later deletes expired rows, and the prompt hash makes answers from an older prompt ineligible as soon as instructions change.
The cache holds user input and model output, so it needs the same deletion paths as any other store of customer data: delete one entry, delete everything for an organisation (a bulk erase for privacy requests, restricted to owners), and cascade on account deletion. Deletion should tombstone first and hard-delete on a schedule, so a mistaken bulk delete can be investigated before the data is gone.
Writes need a gate too. Check each generated answer against the use case's structural and safety rules before making it reusable. At read time, regenerate a small labelled sample of would-be hits and compare it with the cached answer. That sample supplies ongoing precision data and detects drift after traffic, prompts, or models change.
Rolling one out
A semantic cache should earn the right to serve, one step at a time:
From offline test to serving, one tier at a time
Offline: replay real traffic against itself
Take a sample of past requests, hold each one out, and see what the cache would have returned for it. Label a sample of those would-be hits by hand.
outputs: hit rate and wrong-hit rate per use case, per candidate floor
Shadow: look up on every request, serve nothing
The cache fills and every lookup runs, but the provider always answers. Record every would-be hit with its similarity and the entry it matched, and review a sample.
a shadow mode that stores answers but records no would-be hits measures nothing
Serve one tier or one use case
Turn serving on where the measured wrong-hit rate is acceptable, with a thumbs-down path for users and a kill switch that doesn't need a deploy.
gate: wrong-hit rate below the agreed limit for two weeks
Watch it per use case
Hit rate, wrong-hit reports, and savings, split by use case. A blended number hides a use case that never hits and still pays for every lookup.
a use case below a few percent hit rate should bypass the cache
The economics
A hit avoids the generation call. Every request still pays for an embedding and a vector lookup, including misses. Work the numbers before assuming the cache is cheaper.
The following illustrative calculation uses prices published in September 2026: text-embedding-3-small at $0.02 per million input tokens and GPT-5.6 Terra at $2.00 per million input tokens and $12.00 per million output tokens. GPT-5 mini is deprecated, with Terra listed as its replacement. The calculation excludes database, network, and engineering cost.
| work | tokens | calculation | cost |
|---|---|---|---|
| Embed the 20-token request | 20 input | 20 ÷ 1M × $0.02 | $0.0000004 |
| Generate on a miss | 80 input + 120 output | 80 ÷ 1M × $2.00 + 120 ÷ 1M × $12.00 | $0.0016000 |
| Semantic-cache hit | 20 embedding input | embedding only; lookup excluded | $0.0000004 |
| Semantic-cache miss | embedding + generation | $0.0000004 + $0.0016000 | $0.0016004 |
Across one million requests at an assumed 10% hit rate, avoided generation would be worth about $160.00 and embedding every request would cost about $0.40, leaving $159.60 before vector storage and lookup. Ignoring database cost, the token-only break-even hit rate is embedding cost per request ÷ avoided generation cost per hit: $0.0000004 ÷ $0.0016 = 0.00025, or 0.025%. Database work, evaluation, invalidation, and operations raise the real break-even point. A hit can also remove generation latency, which may matter even when token savings are small.
Record hits beside ordinary calls, including the source entry, score, estimated avoided tokens, and actual embedding cost. When computing savings, count uses after the entry was created. The first request paid for the generation that populated the cache.
A checklist
Before building one
- Provider prefix caching is already in use, with stable content first
- The use case is single-shot, short, and repetitive
- A labelled set exists to measure wrong hits
The key
- Tenant on every query, with cross-tenant tests
- Use case, system prompt hash, and embedding model in the key
- A written decision about the generation model
- Request deduplication never crosses tenants
The lookup
- The similarity floor is inside the query
- Floors set per use case from measurements
- Approximate indexes tested for recall under filters
Excluded from semantic response caching
- Agent turns and multi-turn conversations
- Structured outputs that feed code
- Requests where a near-miss is a wrong answer
Lifecycle
- Recency window, retention job, tombstones
- Per-entry and per-organisation deletion
- Quality gate on write, feedback and sampled regeneration on read
Rollout
- Offline replay, then shadow with recorded would-be hits
- Serve one tier or use case at a time, with a kill switch
- Hit rate and savings tracked per use case
Sources
- InstCache, arXiv 2411.13820 (2025); TweakLLM, arXiv 2507.23674 (2025); vCache, arXiv 2502.03771; CacheAttack, arXiv 2601.23088 (2026); MeanCache, arXiv 2403.02694
- LiteLLM, Semantic caching and Caching
- Portkey, Simple and semantic cache; Kong, AI Semantic Cache; Cloudflare, AI Gateway caching
- Redis, Semantic cache and LangCache
- Anthropic, Prompt caching; OpenAI, Prompt caching
- OpenAI, text-embedding-3-small, GPT-5.6 Terra, and model deprecations (checked September 2026)
- pgvector, README: filtering and iterative index scans