embedcache vs Redis vs GPTCache: Caching Embeddings

How embedcache (local embedding plus a SQLite cache) compares with a DIY Redis cache and GPTCache's semantic cache, and what each one actually stores.

The problem

Retrieval pipelines embed the same text over and over. A chunk is embedded at ingest, then again after a re-index, again when a chunking parameter changes, and again when someone re-runs an evaluation. None of those repetitions produces a different vector. Against a hosted embedding API, every one of them is billed, rate-limited and network-bound.

The fix is embedding caching: store the vector for a given input, look it up on the next request, and only run the model on a miss. How much that saves depends entirely on how often identical inputs recur in your pipeline, which is something to measure rather than assume.

embedcache is Skelf’s take on the problem. This post compares it with the two alternatives people most often reach for, plus the one most teams start with.

What embedcache is

embedcache is a Rust library and REST service (GPL-3.0, published on crates.io) that does two things:

  1. Generates embeddings locally. It runs models through the FastEmbed crate on the ONNX runtime. More than twenty models are available, including BGE small/base/large, MiniLM, Nomic and multilingual E5. It is CPU-first; no GPU is required.
  2. Caches them in SQLite. Each vector is stored under a key derived from a hash of the input content plus the model identifier, in a single SQLite file (DB_PATH, default cache.db). Entries are not evicted by default.

The important distinction: embedcache replaces the hosted embedder rather than memoising it. There is no per-token bill and no rate limit, and text never leaves the machine. If you want to keep a hosted model and stop paying for repeats, that is a different pattern, such as LangChain’s CacheBackedEmbeddings or the Redis approach below.

It is also not a vector database. It produces and caches vectors; index and search them in something else (Qdrant, pgvector, LanceDB, or memista for small corpora).

The service exposes three endpoints: POST /v1/embed for an array of strings, POST /v1/process to fetch a URL, chunk it, embed the chunks and cache them, and GET /v1/params to list available models and chunkers. OpenAPI docs are served at /swagger, /redoc, /rapidoc and /scalar. The default chunker splits on whitespace; optional LLM-driven chunkers use Ollama or OpenAI.

What each option is

embedcache bundles the embedder and the cache: local FastEmbed models, a SQLite store keyed by content hash plus model, exact-match lookups.

Redis is the general-purpose cache. The DIY pattern is: hash the text, store the vector as a serialised blob under that key, and decide on eviction and persistence yourself. You bring the embedder, the key derivation and the serialisation. It works, and if you already run Redis it may be the obvious choice.

GPTCache is an open-source Python library from Zilliz for semantic caching of LLM responses. It embeds incoming prompts, looks for similar previous prompts in a vector store, and returns the cached response on a close match. It uses embeddings to find cache hits; its main job is caching model outputs, not serving as a store of embedding vectors.

Custom in-memory dict is what most teams start with. Fine for a prototype; gone on every restart.

The comparison

DimensionembedcacheRedis (DIY)GPTCacheCustom dict
What it cachesEmbedding vectorsWhatever you storeLLM responsesWhatever you store
Embedder bundledYes (FastEmbed, local)NoUses one to match promptsNo
Key strategyContent hash + model idYou derive itPrompt similarityYou derive it
MatchingExactExact (your hash)Similarity thresholdExact (your hash)
StorageOne SQLite fileRedis (in memory, optional persistence)Pluggable storesProcess memory
EvictionNone by defaultYour maxmemory-policyConfigurableNone, until restart
Survives restartYesIf persistence is enabledDepends on storeNo
Multi-language accessREST APINative Redis clientsPythonSame process only
LicenseGPL-3.0Depends on Redis versionOpen sourcen/a
MaturityEarlyWidely usedWidely usedn/a

Competitor details are summarised at the time of writing; check each project’s documentation, particularly Redis licensing, which has changed across versions.

When to use which

Use embedcache when:

  • You want to stop calling a hosted embedding API, not just call it less.
  • A local FastEmbed model is good enough for your retrieval quality.
  • You run on a single node or a small fleet and want one thing to install: the embedder and the cache together.
  • Data residency rules out sending text to a third party.

Use Redis when:

  • You already run Redis and the marginal cost is low.
  • Several services in several languages need to share one cache.
  • You need a model embedcache doesn’t bundle, including a proprietary hosted one.
  • You have decided on eviction and persistence. For embedding workloads that usually means no eviction, or a generous ceiling, plus persistence enabled.

Use GPTCache when:

  • Your problem is repeated LLM calls, not repeated embedding calls.
  • You accept that a similar-but-different prompt may return a cached answer, and you can tune for it.

Use a custom dict when:

  • You are prototyping and the cache is per-process and short-lived.

A worked example (illustrative arithmetic)

This is arithmetic under stated assumptions, not a measurement of embedcache or of any real pipeline.

Assume a pipeline makes 60,000 embedding requests a day (queries plus re-embedded chunks), and assume 95% of those requests hit the cache.

  • Requests that still reach the model: 60,000 × 0.05 = 3,000 a day.
  • Requests served from the cache: 57,000 a day.

If the embedder is a hosted API billed per token, then under that assumption the bill for those requests falls by 95%, because it scales with misses. If the embedder is local, as with embedcache, there is no per-token bill to begin with; the hit rate instead decides how much CPU time the model spends.

Notice what the example proves: the reduction equals the hit rate, by construction. It tells you nothing about your system until you know your hit rate. Two things govern it:

  • Document chunks. For a stable corpus, a content-hash cache with no eviction should approach a 100% hit ratio from the second full re-ingest onward, since unchanged chunks are byte-identical. Measure the hit ratio over a full re-ingest cycle; if it is well below that, something upstream (extraction, chunking, whitespace) is changing the text.
  • Queries. A query cache only hits when the exact string recurs. Some workloads repeat a lot, many repeat very little. Log a week of queries and count exact duplicates before you count on savings.

The semantic-similarity trade-off

embedcache uses exact-match keys. It never returns a vector for different text, but it also treats inputs that differ only in whitespace, punctuation or case as distinct. If you want those to collapse, normalise the text before it reaches the cache, and treat the normalisation as part of your key definition: change it and you should expect a round of misses.

GPTCache’s semantic matching returns a cached result when a new input is close enough to an old one. For LLM responses that can be a reasonable trade. For embeddings it is a correctness problem: an approximate hit during an index build silently stores the wrong vector, and the failure surfaces weeks later as poor relevance rather than as an error.

One caution applies to every content-keyed embedding cache. The key is only correct while the model behind the model id is stable. If a model is upgraded in place without the identifier changing, the cache will serve old vectors against a new index. Put the model version in the key, and make an explicit cache invalidation part of any model upgrade.

A short embedcache eval

# Install the service
cargo install embedcache

# Run it (defaults: 127.0.0.1:8081, cache.db)
embedcache

# Embed a string
curl -X POST localhost:8081/v1/embed \
  -H "Content-Type: application/json" \
  -d '{"model":"BGESmallENV15","text":["hello world"]}'

# Send the same request again: this time it is a cache read

Configuration is by environment variable or .env file: SERVER_HOST, SERVER_PORT, DB_PATH and ENABLED_MODELS, plus LLM_PROVIDER, LLM_MODEL and LLM_BASE_URL if you want the LLM chunkers. GET /v1/params lists the models in your build, and the full request and response schemas are in the OpenAPI docs at /swagger. From Rust, depend on the embedcache crate and call it in-process with no network hop.

  • embedcache repository
  • polymathy, a retrieval pipeline that sends pages to a content processor for chunking and embedding
  • memista, an experimental embedded vector index for small corpora
  • GPTCache, semantic caching for LLM responses
  • Redis, the general-purpose cache