Mohammed Mutahar
← Back to Projects

Semantic Caching Layer

A semantic cache in front of Gemini that matches queries by meaning, with an evaluation harness that sweeps the similarity threshold to keep the operation cost efficient.

Semantic CachingVector SearchFAISSEmbeddingsLLM Cost OptimizationEvaluation & BenchmarkingRedisFastAPILatency OptimizationPython
Semantic Caching Layer

Semantic Cache

A cache that matches meaning instead of strings

A caching layer that sits in front of Gemini and answers repeated questions without calling the model. Incoming queries get embedded, matched against previously answered queries by cosine similarity over a FAISS index, and if the best match clears a threshold, the stored response comes straight back out of Redis.

The hard part was never the caching. It was deciding how similar is similar enough — and then actually measuring it instead of guessing.


Why exact-match caching fails here

Traditional caching keys on the request. Same bytes in, same bytes out, and if a single character differs you've got a miss. This works beautifully for URLs and database queries, and it is nearly worthless for natural language, because nobody asks the same question the same way twice.

What is your refund policy? · How do refunds work? · Can I get my money back? · what's the return policy

Four requests, one answer, four cache misses, four generation calls. From the cache's point of view these have nothing in common. From the user's point of view they're the same question, and paying full price to generate the same answer four times is just waste — in money, and more noticeably in latency, since a cached response returns in milliseconds and a generated one does not.

Embeddings collapse those four into roughly the same region of vector space. That's the whole premise: key the cache on the meaning of the request rather than its literal text, and suddenly the hit rate on real traffic stops being embarrassing.


How it works

Every request follows the same path. The query gets embedded with gemini-embedding-001. That vector goes to a FAISS index holding every query answered so far, which returns the nearest neighbours and their similarity scores. If the top match clears the configured threshold, the corresponding response is pulled from Redis and returned. If it doesn't, the request goes to Gemini, and the new query–response pair is embedded, indexed, and stored for next time.

The two-store split is deliberate. FAISS does one thing well — approximate nearest-neighbour search over vectors — and it is not where you want to keep payloads. Redis holds the actual responses along with TTLs, so expiry is handled by infrastructure that was designed for it rather than bolted onto an index. FAISS answers which stored query is closest; Redis answers what did we say last time.

Every response carries its own diagnostics back: cache_hit, the similarity score that produced the decision, the matched entry_id, and latency_ms. That was the first thing I built and the thing I'd insist on again — a cache that won't tell you why it hit is impossible to tune, and you cannot debug a false hit after the fact if the score that caused it was never recorded.


The threshold is the entire problem

Everything above is about a hundred lines of plumbing. The actual engineering question is one float, and it has genuinely asymmetric consequences.

Set it too high and the cache almost never fires. You've now added an embedding call to every single request and bought nothing with it. Strictly worse than no cache.

Set it too low and the cache serves confidently wrong answers. This is the failure that matters. A false hit isn't a degraded response — it's the answer to a different question, delivered with no indication that anything went wrong. The user asked about international shipping and got the domestic return policy, and nothing in the system flagged it, because from the cache's perspective it did its job perfectly. A miss costs money. A false hit costs trust.

So the two failure modes aren't symmetric and the threshold can't be set by feel. Which is why the sweep exists.

Measuring it

The evaluation harness runs against a dataset of query pairs labelled with whether they're actually equivalent — schema is {query_a, query_b, equivalent} — and sweep_threshold walks the threshold across its range, scoring hit behaviour at every step and writing out a report plus a curve. Point EVAL_DATASET_PATH at your own labelled set and you can tune against your traffic instead of mine, which matters, because the right threshold for a narrow support FAQ and the right threshold for open-domain questions are not the same number.

Two things make the sweep practical rather than theoretical. Embeddings are cached to disk, so the first eval run pays the API cost and every subsequent run and sweep is free — without that, sweeping across a dozen thresholds on a real dataset is expensive enough that you'd do it once and never again. And the whole pipeline runs on stub embeddings offline, so the harness itself can be tested and the plumbing verified with no API key and no network. The test suite has the same property: pytest passes with nothing configured.

There's also a judge gate available in the config — the option to spend a cheap verification call on a borderline candidate rather than trusting the similarity score alone. It trades some of the cache's speed advantage back for protection against exactly the false-hit case above, which is a reasonable deal in a domain where a wrong answer is expensive.

Every knob — threshold, TTL, models, top-k, the judge gate — is exposed in config rather than buried. A cache whose behaviour you can't change without editing source isn't tunable, and this system's entire value depends on being tuned.


What it gets wrong

Cosine similarity is not semantic equivalence, and negation is where that gap becomes dangerous. Is this item refundable? and Is this item non-refundable? sit almost on top of each other in embedding space and have opposite correct answers. Embedding models encode topic strongly and logical polarity weakly, so no threshold setting fixes this — raise it high enough to separate that pair and you've killed the hit rate on everything else. This is the strongest argument for the judge gate, and the clearest limit on where a semantic cache belongs at all.

Some queries should never be cached and the system doesn't know which. Anything whose correct answer depends on who's asking or when — account state, order status, anything time-sensitive — will match a previous user's phrasing and serve their answer. There's no per-user namespacing and no notion of a query being uncacheable, which is fine for a shared FAQ surface and would be a real problem the moment requests carry user-specific context. Namespacing the index per tenant is the obvious fix and the thing I'd build first.

TTL is a blunt instrument for staleness. Cached answers expire on a timer, not when the underlying truth changes. If the refund policy is updated, the cache keeps confidently serving the old one until its TTL runs out. Proper invalidation means knowing which entries a given content change affects, which is a much larger problem than this project takes on.

Generating eval data with an LLM is slightly circular. There's a hand-written fixture committed, and an optional generator for building something larger with Gemini. That's a pragmatic way to get volume, but a dataset of equivalence judgments produced by a model inherits that model's notion of equivalence — and the system under test is trying to judge equivalence. The hand-labelled set is the one I actually trust; the generated one is for pipeline coverage.

The economics depend on hit rate, and hit rate depends on traffic. Every request now pays for an embedding to possibly skip a generation. Embeddings are cheap enough relative to generation that this is a good trade at almost any realistic hit rate, but the break-even is real and it's worth knowing where it sits before deploying.


Stack

Python 3.11+, FastAPI for the service, gemini-embedding-001 for embeddings, FAISS for the vector index, Redis for response storage and TTL, and a self-contained evaluation harness with threshold sweeping, disk-cached embeddings, and offline stub mode. Includes a small frontend and a stats endpoint exposing entry count, index size, and live configuration.

Related Projects