polymathy vs Haystack vs LangChain: Retrieval in Rust
Comparing polymathy with Haystack, LangChain and LlamaIndex: a single-endpoint Rust fetch-and-chunk service against full Python RAG frameworks.
This comparison was rewritten in October 2026. An earlier version described endpoints, answer synthesis, a memista backend and performance figures that polymathy does not have.
The question
Should I use polymathy, Haystack, LangChain, or LlamaIndex for my RAG pipeline?
The honest first answer is that polymathy is not the same kind of thing as the other three, and the comparison only makes sense once that is clear.
The 60-second version: Haystack, LangChain and LlamaIndex are Python frameworks for building whole RAG applications. polymathy is a small Rust web service that does one step of that pipeline: it turns a search query into fetched, chunked page content with its source URL attached. It makes no model call and synthesises no answer. What you do with the chunks is up to you.
What each option is
polymathy is an async Rust web service, published as the
polymathy crate under GPL-3.0. Its public surface is a single
endpoint, GET /v1/search?q={query}. For each request it:
- queries a SearxNG instance you configure (
SEARXNG_URL); - takes the first ten result URLs;
- sends each URL, in parallel, to a content processor you
configure (
PROCESSOR_URL), which extracts the text, chunks it and produces embeddings; - assigns sequential chunk IDs and returns a JSON map of
chunk_id -> [source_url, text].
It bundles no SearxNG instance, no chunker or embedding model, no LLM call, no reranker, no persistent index, no authentication and no UI. A USearch index is instantiated per request but, in v0.2, is not used on the public read path. Every chunk carries its source URL, so attribution is a property of the response shape rather than something a prompt has to remember.
Haystack is deepset’s Python framework for building RAG and search pipelines from components: document converters, retrievers, rankers, readers and generators.
LangChain is a widely used Python framework for LLM applications, including RAG, with a large catalogue of integrations and agent tooling.
LlamaIndex is a Python framework centred on ingesting, indexing and querying data for LLM applications.
The dimensions
| Dimension | polymathy | Haystack | LangChain | LlamaIndex |
|---|---|---|---|---|
| Language | Rust | Python | Python | Python |
| Shape | HTTP service, one endpoint | Library | Library | Library |
| Scope | Search -> fetch -> chunk map | Full pipelines | Full applications | Full pipelines |
| Source of documents | Live web results via SearxNG | Your documents, many loaders | Your documents, many loaders | Your documents, many loaders |
| Chunking and embedding | Delegated to your content processor | Built-in components | Built-in components | Built-in components |
| Persistent index | No | Via document stores | Via vector store integrations | Via vector store integrations |
| Answer generation | No | Yes | Yes | Yes |
| Agents | No | Yes | Yes | Yes |
| Licence | GPL-3.0 | Apache-2.0 | MIT | MIT |
We have not published throughput, memory or start-up measurements for polymathy, and we do not quote any for the frameworks either. If those numbers matter to your decision, measure them on your workload; the fetch step, which depends on remote sites, is likely to dominate.
When to use which
Use polymathy when:
- You already run SearxNG, or are willing to, and want the fetch-and-chunk layer behind a stable HTTP contract.
- You are building a Perplexity-style answer interface and want the retrieval half — fetched, chunked, attributed page content — as a separate Rust service from the generation half.
- You are developing a chunking or embedding service and want a harness that drives it against real-world URLs.
- You want each step inspectable: the chunk map is the intermediate artefact, so bad extraction or bad chunking is visible before it becomes a confidently wrong answer.
Use Haystack when:
- You are building a production RAG or search pipeline in Python over your own documents.
- You want ready-made converters for PDFs, office documents and web pages, and a choice of retrieval strategies.
Use LangChain when:
- You are building a broader LLM application in which retrieval is one component among many.
- You want agent capabilities and a large catalogue of integrations.
Use LlamaIndex when:
- Your problem is mainly ingesting and indexing a varied body of data and querying it.
The polymathy design
The key design choice is scope. A pipeline that goes from query to generated answer in one call cannot be inspected in the middle, and the middle is where retrieval fails: the extraction picked up navigation furniture, the chunking split a table, the source was irrelevant. polymathy returns that middle as its product.
Configuration is a handful of environment variables, read from a
.env file or the process environment: SEARXNG_URL and
PROCESSOR_URL (both required), and SERVER_HOST and
SERVER_PORT (defaulting to 127.0.0.1 and 8080). Running it
from source and querying it looks like this:
git clone https://github.com/skelfresearch/polymathy.git
cd polymathy
cp sample.env .env # set SEARXNG_URL and PROCESSOR_URL
cargo run --release
curl "http://localhost:8080/v1/search?q=remote+work+policy"
The response is the chunk map:
{
"0": ["https://example.com/handbook/remote", "Employees may work remotely up to..."],
"1": ["https://example.com/handbook/equipment", "Equipment for home working is..."]
}
An OpenAPI specification generated from the Rust types is served
at /openapi.json, alongside Swagger and ReDoc interfaces.
Within our portfolio, polymathy does not depend on embedcache or
memista. Chunking and embedding are delegated to whatever
PROCESSOR_URL points at, and polymathy keeps its own USearch
index. embedcache (local embedding generation with an exact-match
cache) and memista (SQLite plus a USearch HNSW index) are
separable components that a team could put either side of it.
That separation is deliberate: the interfaces between them are
ordinary HTTP, not a shared internal abstraction.
A concrete example: an internal-docs answer engine
Suppose you want a Q&A interface over an internal wiki.
With polymathy, you point a SearxNG instance at your internal
search indices, run a content processor that chunks and embeds
pages, and run polymathy between them. Your application calls
/v1/search, receives attributed chunks, and passes them to
whichever model you choose to write the answer. You own the
generation step, the prompt and any persistent index.
With Haystack, LangChain or LlamaIndex, you load the wiki pages directly with a document loader, chunk and embed them into a vector store, and build a pipeline that retrieves and generates in one place. Less to operate as separate services, and much more functionality out of the box.
Both work. The first suits teams who want to own each stage behind HTTP; the second suits teams who want a working system quickly in Python.
When polymathy is the WRONG answer
- You want a RAG system, not a component. polymathy does not generate answers, keep an index or serve a UI. The Python frameworks do far more.
- Your documents are not reachable through search. polymathy’s input is SearxNG results; its recall is capped by the backends configured behind SearxNG.
- You need document loaders. polymathy fetches URLs and hands them to your processor. Parsing PDFs or office files is the processor’s job, or a framework’s.
- You need production hardening out of the box. There is no authentication, rate limiting or multi-tenant isolation.
- You are not prepared to run three services. SearxNG, a content processor and polymathy are each yours to operate.
What to read next
- polymathy repository
- embedcache — local embedding generation with an exact-match cache
- memista — SQLite-backed vector search
- Haystack
- LangChain
- LlamaIndex