0%

0000000

0x00

Qdrant and BGE-M3: Hybrid Search Without a Second Model

Most RAG problems are retrieval problems. Dense and sparse result lists crossing into one fused list: one model, two vectors, fused.

BGE-M3 produces dense and sparse vectors from one pass, so Qdrant can run hybrid search for production RAG without a second embedding model: prefetch both, fuse with RRF, filter early.

TL;DR: Most RAG problems are retrieval problems wearing an LLM costume. BGE-M3 produces a dense vector and a sparse lexical vector from the same pass over a text, so Qdrant can run hybrid search for production RAG without a second embedding model: store both as named vectors, query them together with prefetch, fuse the results with Reciprocal Rank Fusion, and filter on indexed payload fields so only matching points reach the fusion step. Then measure retrieval on its own, before any model writes an answer.

Why Are Most RAG Problems Retrieval Problems?

Most RAG problems are retrieval problems because a language model can only answer from the passages it is handed. When the right passage is not in the top results, a better prompt or a bigger model produces a more fluent wrong answer, and the failure looks like a model problem.

I build RAG pipelines on Qdrant with BGE-M3 embeddings for client projects, with self-hosted inference on Ollama and Qwen3. Self-hosted retrieval follows the same logic as what an AWS to self-hosted migration really costs: when documents have to stay on hardware you control, hosted embedding APIs and managed vector search are off the table, and there is no vendor to blame when retrieval misses.

The decisions that matter most sit before generation: which vectors to store, how to combine them, and how to filter. The rest of this guide covers those, with every claim taken from Qdrant's and BGE-M3's own documentation.

How Does BGE-M3 Give Qdrant Dense and Sparse Vectors From One Model?

BGE-M3 gives Qdrant both vector types from one model because it was trained to do three kinds of retrieval at once. Its model card lists "Dense retrieval", "Sparse retrieval (lexical matching)" and "Multi-vector retrieval", and says it "can simultaneously perform the three common retrieval functionalities" on input "from short sentences to long documents of up to 8192 tokens."

The upshot: one embedding service instead of two. Hybrid setups often pair a dense model with a separate BM25 retriever; Qdrant's own fusion docs use exactly that pairing, noting that "a dense retriever may dominate on natural-language queries, while BM25 may win on identifier-heavy ones." With BGE-M3, the reference implementation in FlagEmbedding returns both from a single call:

from FlagEmbedding import BGEM3FlagModel

model = BGEM3FlagModel('BAAI/bge-m3', use_fp16=True)
output = model.encode(chunks, return_dense=True, return_sparse=True)

dense = output['dense_vecs'].tolist() # one 1024-dimension vector per chunk
sparse = output['lexical_weights']    # one {token_id: weight} map per chunk

The model card gives the rest of the shape: 1024-dimension dense embeddings, a maximum input of 8192 tokens and support for over 100 languages. The long input limit changes how chunking should be decided. BGE-M3 can embed a whole section in one vector, so chunk size becomes a retrieval-precision choice rather than a model constraint: smaller chunks put one idea in each vector, and larger chunks keep context together at the cost of diluting it.

How Do You Run Qdrant Hybrid Search With BGE-M3 for Production RAG?

Store the dense and sparse outputs as two named vectors on each point, then query both through the Query API's prefetch and fuse the results. Qdrant's documentation puts the reason in one line: combining them gives "semantic understanding from dense vectors and precise word matching from sparse vectors."

The collection declares both vector types:

from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333")

client.create_collection(
    collection_name="docs",
    vectors_config={"dense": models.VectorParams(size=1024, distance=models.Distance.COSINE)},
    sparse_vectors_config={"sparse": models.SparseVectorParams()},
)

BGE-M3's lexical weights map token ids to weights, which converts directly into Qdrant's sparse format:

def to_sparse(weights):
    return models.SparseVector(
        indices=[int(token) for token in weights],
        values=[float(weight) for weight in weights.values()],
    )

The query runs both searches, then fuses them:

hits = client.query_points(
    collection_name="docs",
    prefetch=[
        models.Prefetch(query=dense_query, using="dense", limit=20),
        models.Prefetch(query=to_sparse(sparse_query), using="sparse", limit=20),
    ],
    query=models.FusionQuery(fusion=models.Fusion.RRF),
    limit=5,
)

Qdrant's docs describe exactly what happens: with at least one prefetch, Qdrant will "Perform the prefetch query (or queries)", then "Apply the main query over the results of its prefetch(es)." Here the main query is Reciprocal Rank Fusion, which ranks by position in each list rather than by raw score, so the dense and sparse scores never have to be made comparable. Qdrant's docs make the same point: dense and sparse scores "live on different scales", and "RRF sidesteps this by using ranks."

Qdrant offers a second fusion method. Distribution-Based Score Fusion "keeps the raw scores from each query but normalizes their distributions before combining." RRF has been available since Qdrant 1.10.0 and DBSF since 1.11.0. Qdrant's own guidance is to start with RRF, "the safe default", move to weighted RRF (available since 1.17.0) once an evaluation set exists to tune the weights on, and reserve DBSF for retrievers whose raw scores you trust.

How Do Payload Filters Change a Hybrid Qdrant Query?

Payload filters change a hybrid Qdrant query by restricting which points the dense and sparse searches can return before fusion ranks them. One filter on the main query covers the common case: on Qdrant 1.19.2, a top-level query_filter was applied to the prefetch candidates as well, so filtered results still filled the limit. A filter inside a single Prefetch scopes only that sub-search.

by_type = models.Filter(
    must=[models.FieldCondition(key="doc_type", match=models.MatchValue(value="protocol"))]
)

hits = client.query_points(
    collection_name="docs",
    prefetch=[
        models.Prefetch(query=dense_query, using="dense", limit=20),
        models.Prefetch(query=to_sparse(sparse_query), using="sparse", limit=20),
    ],
    query=models.FusionQuery(fusion=models.Fusion.RRF),
    query_filter=by_type,
    limit=5,
)

Filters are only fast with an index behind them. Qdrant's documentation is direct: "For performant filtering, create payload indexes for the fields you plan to filter on", and "For best results, create payload indexes before ingesting data." A payload index is "built for a specific field and type", with keyword for exact matches, integer and float for ranges, datetime for date ranges and uuid for identifiers:

client.create_payload_index(
    collection_name="docs",
    field_name="doc_type",
    field_schema=models.PayloadSchemaType.KEYWORD,
)

With the index in place, Qdrant combines it with the vector index into what its docs call a "filterable HNSW Index", where "extra edges allow you to efficiently search for nearby vectors using the HNSW index and apply filters as you search in the graph." Decide the filter fields at design time: document type, source, date and, in a multi-tenant system, the tenant.

How Do You Check Retrieval Quality Separately From the Answers?

You check retrieval quality separately by testing the retriever on its own, with a fixed set of real questions and the passages that should come back for each, before any model generates text. If the right passage is not in the top results, no answer evaluation will tell you why the answer was wrong.

A retrieval check needs three things:

  1. A question set from real users or real documents, each question paired with the chunk or chunks that answer it.
  2. One number per configuration: how often the right chunk appears in the top results.
  3. Runs that change one thing at a time: dense only against hybrid, RRF against DBSF, one chunk size against another, the same filters each time.

Run the same set after every change to chunking, the embedding model or the fusion method, and only then evaluate the generated answers. Separating the two turns "the bot gave a wrong answer" into either "retrieval missed it" or "the model misread it", and each has a different fix.

Limitations

  • This is a design guide built on Qdrant's and BGE-M3's own documentation, for the stack I build RAG pipelines on. Client work is under NDA, so nothing from client projects is quoted here.
  • Every Qdrant snippet above was run on 2026-10-06 against Qdrant 1.19.2 with qdrant-client 1.19.1: dense and sparse prefetch with RRF, the filtered query and DBSF all returned results. The FlagEmbedding call is the one from the BGE-M3 reference README.
  • BGE-M3's third output, multi-vector ColBERT-style retrieval, is not covered here.
  • Fusion and chunk size have no universal best setting. The evaluation set is what decides them for a given corpus.

Hybrid search earns its complexity only when you can measure what it changed. Are you running hybrid search in production, and did it earn its place over dense retrieval alone? More on running AI on your own hardware lives in the Self-Hosted AI category.

References