Skip to content

Architecture Article

Semantic Caching for LLM Calls

Learn when a semantic cache can safely reuse an LLM answer, how false hits happen, and how tenant boundaries, evaluation, and freshness limit the risk.

Published 20 Aug 2026Updated 1 Oct 20268 min read
LLM · Semantic Caching · pgvector · Cost Optimization · AI Gateway
Questions arranged around a similarity boundary, with one near-looking question rejected as unsafe.
On this page (12)

A semantic cache answers a new request with an answer created for an earlier, different request. It decides that the two requests are close enough in meaning. A correct match avoids the model call. A false match returns a polished answer to the wrong question, which can be harder to notice than an error.

Use the provider's exact-prefix prompt cache where its contract fits; it reuses prompt computation while the model still produces a new answer. Use a semantic response cache only for narrow, repetitive, short tasks, inside one tenant boundary, after measuring false hits. Do not apply it to an agent's coordinator turns or to a whole multi-turn conversation.

How to read the diagrams

  • The request and the cache's records
  • A correct reuse: the stored answer fits
  • A false hit: the stored answer is for a different question
  • A provider call: the miss path, where tokens cost money
  • Tenant and policy boundaries

Three kinds of reuse

Exact response cache

Matches
A byte-identical request inside the same complete cache key
Saves
The whole call
Risk
Low only when tenant, instructions, model policy, permissions, and freshness are also part of eligibility
Example
Cloudflare AI Gateway caches identical requests only

Provider prompt cache

Matches
An identical prefix of the prompt
Saves
Input processing: in September 2026, GPT-5.6 Terra cached input is $0.20/M versus $2.00/M uncached; Anthropic rates vary by model and cache duration
Risk
It does not reuse the prior final answer; provider eligibility and expiry rules still apply
Example
Supported OpenAI and Anthropic models

Semantic cache

Matches
A different request with a close embedding
Saves
The whole call, output included
Risk
Serving the answer to a different question
Example
GPTCache, Redis LangCache, gateway plugins

Exact and provider-prefix caches reuse work for identical bytes. Identical input is not enough by itself: an exact response can still be stale or belong to another tenant or prompt version, so it needs the same policy boundaries. The semantic cache crosses an additional boundary and decides that different text deserves the same answer. The rest of this chapter is about making that decision visible and measurable.

How text becomes a similarity score

An embedding turns text into a vector, a list of numbers that places the text on a model's learned map. Nearby directions on that map often express related ideas. Cosine similarity compares the angle between two vectors: 1 means the directions are identical, lower values mean they point farther apart. The score describes the embedding model's geometry; it does not prove that two requests have the same correct answer.

A nearest-neighbour search asks which stored vectors sit closest to the new vector. A threshold is the minimum score the application will accept as a cache hit. The map analogy stops here: names, dates, negation, and product boundaries can change the correct answer without moving the vector very far.

The following scores are illustrative. The cached request is “Can you move my demo to Thursday?” and this example uses a threshold of 0.92.

questions compared with one cached requestillustrative cosine similarities · threshold 0.92
new request
Please reschedule my demo for Thursday
similarity
0.96
cache decision
hit
should reuse?
yes
new request
Move my onboarding call to Thursday
similarity
0.94
cache decision
hit
should reuse?
no: different appointment
new request
Can you move Jordan's demo to Thursday?
similarity
0.92
cache decision
hit
should reuse?
no: different person
new request
Can you move my demo to Friday?
similarity
0.91
cache decision
miss
should reuse?
no: different date
new request
Can you cancel my demo on Thursday?
similarity
0.89
cache decision
miss
should reuse?
no: opposite action
new request
What will the weather be on Thursday?
similarity
0.41
cache decision
miss
should reuse?
no: unrelated
The first hit is useful. The next two are false hits: the words are close while the business action is different.

What one lookup does

The running example is a use case that suggests a short reply to an inbound customer message. Here is one lookup, with the filters that decide which stored answers are even eligible:

suggest_reply · one cache lookup

  1. Embed the request

    The user's message is turned into a vector with the same embedding model used for the stored entries.

    illustrative: text-embedding-3-small · 1,536 dimensions

  2. Filter to eligible entries

    Only entries inside the same tenant and policy boundary can match, and only while they are fresh.

    illustrative: org_42 · suggest_reply · prompt h:9c2e… · policy v4 · not expired

  3. Find neighbours above the floor

    The threshold is inside the query. The closest entry is still a miss when its score is too low.

    illustrative: cosine similarity ≥ 0.92 · top 5

  4. Apply the serving policy

    Similarity is necessary but not sufficient. The entry must also have passed quality checks when it was written.

    quality ok · not sampled for regeneration today

  5. Serve the stored answer

    The request does not reach the generation model. Record the source entry and score so the decision can be audited.

    illustrative: hit cache_17 · score 0.94

In SQL, with pgvector, the lookup looks like this:

Sketch: a scoped nearest-neighbour lookup with the floor inside the query
SELECT id, response, 1 - (embedding <=> $1) AS similarity
FROM llm_cache
WHERE org_id           = $2        -- tenant: never optional
  AND use_case         = $3
  AND prompt_hash      = $4        -- a new system prompt starts a new cache
  AND embedding_model  = $5
  AND policy_version   = $6
  AND locale           = $7
  AND quality_status   = 'approved'
  AND deleted_at IS NULL
  AND fresh_until      > now()
  AND (embedding <=> $1) <= 1 - $8 -- $8 = similarity floor
ORDER BY embedding <=> $1
LIMIT 5;

The distance condition belongs in the WHERE clause. Without it, the query returns the nearest neighbour even when the nearest neighbour is not similar at all, and correctness then depends on every caller remembering to check the score.

The threshold is the product

Similarity scores do not know the cost of a wrong answer. That makes the threshold a product decision rather than a model default. Here is the running request against four stored requests:

nearest neighbours for: “Can you move my demo to Thursday?”illustrative similarity scores
stored requestsimilaritysame answer?
Please reschedule my demo for Thursday0.96yes: same intent, same day
Could you move my demo to Friday?0.93no: a different day
Can you cancel my demo on Thursday?0.91no: the opposite request
What time is my demo on Thursday?0.88no, and below the floor

A floor of 0.90 serves the first row, which is right, but also the second and third, which are wrong. Raising the floor to 0.935 fixes this tiny example but cannot prove the next dataset will behave the same way. Small words that change the answer, such as Friday, cancel, and not, may move an embedding less than expected.

Measure the trade-off on labelled requests from the actual use case. Precision asks: of the requests served from cache, how many were correct? Recall asks: of all requests that could safely have reused an answer, how many did the cache find? A higher threshold normally raises precision and lowers recall.

100 labelled requests at three candidate thresholdsillustrative sample · 50 requests have a safe reusable answer
thresholdcorrect hitswrong hitssafe matches missedprecisionrecall
0.884212877.8%84%
0.922942187.9%58%
0.961213892.3%24%
The strictest threshold still makes one wrong reuse and misses most safe opportunities. The acceptable row depends on the harm of a false hit.

Published measurements show why a threshold needs local evaluation. In an appendix experiment on English WildChat samples, the InstCache authors used GPTCache with its Albert-small embedding model and used DeepSeek V3 to judge whether matched instructions would have similar answers:

GPTCache on WildChat conversations (InstCache, 2025)

Share of requests served from cache

Raising the floor lowers the hit rate, as expected.

Share of those hits that were wrong

It doesn't lower the error rate: about a third of hits were mismatches at both floors.

y: percent

TweakLLM measured GPTCache on a question-pairs dataset. At threshold 0.70, its reported precision was about 0.90; with an ALBERT re-ranker at threshold 0.97, precision reached 0.97 while recall fell to about 0.20. vCache compared its dynamic threshold with the static-threshold and other baselines in its own workloads, reporting up to 12.5× more cache hits and up to 26× fewer errors. Those ratios are comparisons inside that experiment, not production guarantees for another embedding model or dataset.

Set the floor per use case. Classification into a small fixed set may tolerate a trade-off that is unacceptable for writing a customer reply. Build should-match and should-not-match pairs, then compute precision and recall at each candidate floor. If the measured false-hit cost is still too high, bypass semantic reuse for that use case.

A product can scope semantic-cache eligibility to a tenant and use case, then evaluate candidate thresholds in shadow mode before serving reviewed matches. The example is a rollout pattern, not a statement about a particular product's plans or active controls.

The cache key is a security boundary

Every filter in the lookup above is there for correctness, and the tenant filter is also there for security. A cache that matches across tenants will eventually serve one customer's data to another, because a similar question from a different company is still a similar question.

Off-the-shelf tools are explicit about where their defaults draw the line. LiteLLM's documentation warns that end users behind one virtual key share a semantic bucket by default, which allows one user's response, including tool calls, to match another user's similar prompt. Portkey's semantic matching ignores the system prompt, so two use cases with different instructions can answer each other's requests; its semantic-cache feature is listed for Enterprise plans as of September 2026. Redis recommends "hard metadata boundaries (tenant, locale, model version, safety flags)". CacheAttack showed that an attacker who can write to a shared semantic cache can craft entries that hijack other users' responses, reporting an 86% attack hit rate in its experiments. Locality makes a semantic cache useful, but it does not provide the collision resistance of a cryptographic key.

One illustrative row makes the boundary concrete:

semantic_cache
field
org_id
example
org_42
if omitted from eligibility
another customer's answer can cross the tenant boundary
field
use_case
example
suggest_reply
if omitted from eligibility
a classifier result can serve a writing request
field
prompt_hash
example
9c2e…
if omitted from eligibility
old instructions keep answering after a prompt change
field
embedding_model
example
text-embedding-3-small
if omitted from eligibility
scores compare vectors from incompatible spaces
field
policy_version
example
4
if omitted from eligibility
an entry approved under an older safety or quality rule remains eligible
field
locale
example
en-GB
if omitted from eligibility
the answer can use the wrong language or formatting convention
field
fresh_until
example
illustrative: day 10
if omitted from eligibility
date-sensitive answers can live forever
field
generation_model
example
recorded, policy decides whether it filters
if omitted from eligibility
a model upgrade may keep serving older-model answers without an explicit decision
field
request · embedding · response
example
encrypted payload + vector + answer
if omitted from eligibility
there is nothing to compare or return; these are stored data, not boundary fields
field
quality status · deleted_at
example
approved · NULL
if omitted from eligibility
rejected or deleted entries can be served
Tenant is the security boundary. The remaining fields prevent answers from crossing policy, language, model, and freshness boundaries.

Database row-level security can reinforce tenant isolation. Tests should attempt a cross-tenant lookup and prove it returns no row. Whether the generation model is part of eligibility is an explicit product decision: including it drops all old-model hits after an upgrade; omitting it saves more calls but keeps serving older-model answers until another boundary changes.

Deduplicating concurrent identical misses (so ten simultaneous copies of the same request make one provider call) is worth doing, and it has the same rule: the deduplication key includes the tenant, so it never merges requests from two organisations.

Do not reuse agent turns as semantic-cache answers

The strongest warning in the published material concerns agents. LiteLLM's docs explain that in an agent loop "consecutive turns are nearly identical text and their embeddings are ~0.99 similar. At any practical threshold, every turn matches the previous turn's cached entry … which typically shows up as an agent repeating the same tool call over and over. Raising similarity_threshold does not reliably fix this."

The reason is structural. In a multi-turn conversation, the new information is a small addition at the end of a long shared history, and an embedding of the whole history is dominated by the shared part. The cache sees "the same conversation" and returns the previous step's answer.

Structured outputs that feed code need the same caution. A semantically close request can require a different identifier, enum member, or tool argument while still producing a high embedding score. Reusing the old JSON can pass schema validation and execute the wrong action, so these calls bypass response reuse.

Keep semantic response reuse to independently answerable requests. An agent can still use exact caches for deterministic tool results and a provider's prefix cache for stable instructions and history. Those mechanisms do not substitute a previous coordinator decision merely because the new turn embeds nearby.

pgvector is a PostgreSQL extension that stores vectors and calculates their distance. An exact search compares the query vector with every eligible row. An approximate nearest-neighbour index searches a smaller part of the space to respond faster, accepting that it may miss a true neighbour. HNSW, short for hierarchical navigable small world, is one such graph-based index; IVFFlat is another.

When tenant and use-case filters leave a small candidate set, a B-tree index on those fields can narrow the rows before an exact distance calculation. Exact search never loses recall because of an approximate index, so start there and measure before adding HNSW.

Approximate indexes become useful when the eligible set is large, and they have a sharp edge with filters. The pgvector README explains that filtering is applied after an approximate scan. If the scan visits 40 candidates and a filter keeps 10% of the overall rows, it keeps about four candidates on average. Iterative scans can search farther when filters remove results, while partial indexes or partitioning can narrow the indexed population. Measure recall by comparing approximate results with an exact scan on the same labelled queries.

Freshness, deletion, and what gets stored

Stored answers age. A time to live, or TTL, is how long an entry stays eligible. The illustrative ten-day window in the lookup is an example, not a production setting. A scheduled retention job later deletes expired rows, and the prompt hash makes answers from an older prompt ineligible as soon as instructions change.

The cache holds user input and model output, so it needs the same deletion paths as any other store of customer data: delete one entry, delete everything for an organisation (a bulk erase for privacy requests, restricted to owners), and cascade on account deletion. Deletion should tombstone first and hard-delete on a schedule, so a mistaken bulk delete can be investigated before the data is gone.

Writes need a gate too. Check each generated answer against the use case's structural and safety rules before making it reusable. At read time, regenerate a small labelled sample of would-be hits and compare it with the cached answer. That sample supplies ongoing precision data and detects drift after traffic, prompts, or models change.

Rolling one out

A semantic cache should earn the right to serve, one step at a time:

From offline test to serving, one tier at a time

  1. Offline: replay real traffic against itself

    Take a sample of past requests, hold each one out, and see what the cache would have returned for it. Label a sample of those would-be hits by hand.

    outputs: hit rate and wrong-hit rate per use case, per candidate floor

  2. Shadow: look up on every request, serve nothing

    The cache fills and every lookup runs, but the provider always answers. Record every would-be hit with its similarity and the entry it matched, and review a sample.

    a shadow mode that stores answers but records no would-be hits measures nothing

  3. Serve one tier or one use case

    Turn serving on where the measured wrong-hit rate is acceptable, with a thumbs-down path for users and a kill switch that doesn't need a deploy.

    gate: wrong-hit rate below the agreed limit for two weeks

  4. Watch it per use case

    Hit rate, wrong-hit reports, and savings, split by use case. A blended number hides a use case that never hits and still pays for every lookup.

    a use case below a few percent hit rate should bypass the cache

The economics

A hit avoids the generation call. Every request still pays for an embedding and a vector lookup, including misses. Work the numbers before assuming the cache is cheaper.

The following illustrative calculation uses prices published in September 2026: text-embedding-3-small at $0.02 per million input tokens and GPT-5.6 Terra at $2.00 per million input tokens and $12.00 per million output tokens. GPT-5 mini is deprecated, with Terra listed as its replacement. The calculation excludes database, network, and engineering cost.

cost of one short reply request
worktokenscalculationcost
Embed the 20-token request20 input20 ÷ 1M × $0.02$0.0000004
Generate on a miss80 input + 120 output80 ÷ 1M × $2.00 + 120 ÷ 1M × $12.00$0.0016000
Semantic-cache hit20 embedding inputembedding only; lookup excluded$0.0000004
Semantic-cache missembedding + generation$0.0000004 + $0.0016000$0.0016004
At these illustrative token counts, one hit saves about $0.0015996 before database cost. One miss costs slightly more than calling the model directly.

Across one million requests at an assumed 10% hit rate, avoided generation would be worth about $160.00 and embedding every request would cost about $0.40, leaving $159.60 before vector storage and lookup. Ignoring database cost, the token-only break-even hit rate is embedding cost per request ÷ avoided generation cost per hit: $0.0000004 ÷ $0.0016 = 0.00025, or 0.025%. Database work, evaluation, invalidation, and operations raise the real break-even point. A hit can also remove generation latency, which may matter even when token savings are small.

Record hits beside ordinary calls, including the source entry, score, estimated avoided tokens, and actual embedding cost. When computing savings, count uses after the entry was created. The first request paid for the generation that populated the cache.

A checklist

Before building one

  • Provider prefix caching is already in use, with stable content first
  • The use case is single-shot, short, and repetitive
  • A labelled set exists to measure wrong hits

The key

  • Tenant on every query, with cross-tenant tests
  • Use case, system prompt hash, and embedding model in the key
  • A written decision about the generation model
  • Request deduplication never crosses tenants

The lookup

  • The similarity floor is inside the query
  • Floors set per use case from measurements
  • Approximate indexes tested for recall under filters

Excluded from semantic response caching

  • Agent turns and multi-turn conversations
  • Structured outputs that feed code
  • Requests where a near-miss is a wrong answer

Lifecycle

  • Recency window, retention job, tombstones
  • Per-entry and per-organisation deletion
  • Quality gate on write, feedback and sampled regeneration on read

Rollout

  • Offline replay, then shadow with recorded would-be hits
  • Serve one tier or use case at a time, with a kill switch
  • Hit rate and savings tracked per use case

Sources