Production-ready

Hybrid Search for Elasticsearch

BM25 lexical search and semantic retrieval merged with Reciprocal Rank Fusion in a single Elasticsearch index. Graceful fallback when inference endpoints aren't available. Deploy it now on any Elasticsearch cluster.

The problem it solves

Most "hybrid search" setups run Elasticsearch (or another search engine) for lexical matching and a separate vector database — Pinecone, Qdrant, Weaviate — for semantic matching. That's a second write path to keep in sync, a second system that can silently drift out of consistency with the first, and a multi-hop query stitched together by hand in application code that gets slow and flaky under load, right when search matters most.

Elasticsearch 8.15+ (and Elastic Cloud/Serverless) can do the semantic half itself — a semantic_text field type that computes embeddings in-cluster via a built-in inference endpoint, no external vector store or embedding API required. This kit is that pattern, extracted into a package you can drop into your own project.

What's included

  • A reusable hybrid_search package (no Flask dependency) — index setup with the right semantic_text mapping, and a query builder that runs BM25 + semantic as one RRF-merged retriever query.
  • Graceful fallback, in real code — the same function that builds the hybrid query automatically returns a plain lexical query when a cluster has no inference endpoint, instead of raising. Proven by unit tests, not just described in prose.
  • A minimal example Flask app wiring the package together end to end, with sample documents chosen so you can see semantic retrieval surface a result that shares none of the query's exact words.
  • 13 unit tests (mocked Elasticsearch client, no cluster required) pinning down the RRF query structure and the fallback branch.
  • A README covering setup and the pitfalls that actually bite — inference-endpoint configuration, ELSER's RAM overhead, indexing lag under bulk load, and why running two separate systems costs more than it looks like up front.

Who it's for

Developers who already run Elasticsearch and have hit (or are about to hit) the two-system hybrid-search problem — bolting a vector database onto an existing search stack instead of turning on a feature the cluster already has. If you haven't started yet, this is the pattern to build toward instead of reaching for a second database by default.

Ask about the kit →

Checkout is opening shortly. Email us and we'll get you a copy.


Deploy on your own infrastructure. Full control, zero vendor lock-in.