What happened?
Description
I am trying to understand how the sparse vectors are generated and scored when using HybridSearchVectorStore with Qdrant and: view code..
I am observing a significant mismatch between the query and the documents returned by the sparse vector search.
For example, after ingesting documents containing content about ASP.NET Core, a completely unrelated query such as:
"¿Qué es Colombia?"
can return chunks related to ASP.NET Core from the sparse vector store.
I would like to understand how this can happen at the vector/scoring level.
Specifically:
How does Qdrant/bm25 transform the document text into a sparse vector during ingestion?
How is the query transformed into a sparse vector?
Which terms and weights are stored in the sparse vector?
How is the similarity score between the query sparse vector and document sparse vector calculated?
Why can a document with apparently no relevant terms to the query obtain a significant non-zero score?
Does Ragbits apply any normalization or transformation to the BM25 score before returning the result?
Does HybridSearchVectorStore apply a score threshold to sparse results before combining them with dense results?
I am particularly interested in understanding the complete path:
Document
-> BM25 sparse vector
-> Qdrant indexed sparse vector
Query
-> BM25 sparse query vector
-> Qdrant sparse search
-> similarity score
-> HybridSearchVectorStore
The main concern is not that BM25 returns a non-zero score, but understanding why documents that appear completely unrelated to the query can obtain sufficiently high scores to be returned as relevant results.
Could you clarify whether this behavior is expected from the Qdrant/bm25 implementation, or whether Ragbits performs any additional processing that could explain this behavior?
Reference:
https://qdrant.tech/articles/sparse-vectors/
How can we reproduce it?
from qdrant_client import AsyncQdrantClient
from ragbits.core.vector_stores.qdrant import QdrantVectorStore
from ragbits.core.vector_stores.base import VectorStoreOptions
from ragbits.document_search import DocumentSearch, DocumentSearchOptions
from ragbits.core.embeddings.dense import LiteLLMEmbedder
from ragbits.core.embeddings.dense.fastembed import FastEmbedEmbedder
from ragbits.core.embeddings.sparse.fastembed import FastEmbedSparseEmbedder
from ragbits.core.vector_stores.hybrid import HybridSearchVectorStore
from ragbits.core.llms import LiteLLM
from ragbits.document_search import DocumentSearch
from ragbits.document_search.ingestion.parsers import DocumentParserRouter
from ragbits.document_search.documents.document import DocumentType
from ragbits.document_search.retrieval.rephrasers import LLMQueryRephraser
from ragbits.document_search.ingestion.parsers.docling import DoclingDocumentParser
from ragbits.core.embeddings.dense.litellm import LiteLLMEmbedderOptions
from ragbits_rag.config import config
from docling_core.transforms.chunker.hybrid_chunker import HybridChunker
def get_llm():
return LiteLLM(
model_name=f"huggingface/{config.llm_model}",
api_base=config.llm_base_url,
api_key=config.openai_api_key,
use_structured_output=True,
)
async def get_vector_store():
store_prefix = "ragbits-rag"
qdrant_client = AsyncQdrantClient(config.qdrant_host)
dense_embedder = LiteLLMEmbedder(
model_name=f"openai/{config.embedding_model}",
base_url=config.embedding_base_url,
default_options=LiteLLMEmbedderOptions(encoding_format="float")
)
sparse_embedder = FastEmbedSparseEmbedder(
model_name="Qdrant/bm25",
use_gpu=False,
)
sparse_store = QdrantVectorStore(
client=qdrant_client,
embedder=sparse_embedder,
index_name=f"{store_prefix}-sparse",
)
dense_store = QdrantVectorStore(client=qdrant_client, embedder=dense_embedder, index_name=store_prefix + "-dense")
return HybridSearchVectorStore(
dense_store,
sparse_store,
)
async def get_document_search():
llm = get_llm()
vector_store = await get_vector_store()
chunker = HybridChunker(
max_tokens=256,
merge_peers=True,
)
# Configure document parsers
parser_router = DocumentParserRouter({
DocumentType.PDF: DoclingDocumentParser(chunker=chunker),
DocumentType.DOCX: DoclingDocumentParser(chunker=chunker),
DocumentType.HTML: DoclingDocumentParser(chunker=chunker),
DocumentType.TXT: DoclingDocumentParser(chunker=chunker),
})
document_search = DocumentSearch(
vector_store=vector_store,
query_rephraser=LLMQueryRephraser(llm=llm),
parser_router=parser_router,
enricher_router=None
)
return document_search
uv run ragbits document-search --factory-path ragbits_rag.components:get_document_search ingest "web://https://www.webcolegios.com/file/51faaa.pdf"
uv run ragbits document-search --factory-path ragbits_rag.components:get_document_search ingest "web://https://pdfobject.com/pdf/sample.pdf"
uv run ragbits document-search --factory-path ragbits_rag.components:get_document_search search "que es colombia?"
import argparse
import asyncio
from ragbits_rag.components import get_document_search
from ragbits.core.vector_stores.base import VectorStoreOptions
from ragbits.document_search import DocumentSearchOptions
async def main(query: str, k: int, threshold: float | None):
document_search = await get_document_search()
options = DocumentSearchOptions(
vector_store_options=VectorStoreOptions(
k=k,
score_threshold=threshold,
)
)
results = await document_search.search(
query,
options=options,
)
print("#" * 75)
for result in results:
print("+" * 75)
print(result)
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument("query")
parser.add_argument("--k", type=int, default=10)
parser.add_argument("--threshold", type=float)
args = parser.parse_args()
asyncio.run(main(
args.query,
args.k,
args.threshold,
))
uv run python src/ragbits_rag/query_rag.py "que es colombia?" --k 10 --threshold 0.55
Relevant log output
What happened?
Description
I am trying to understand how the sparse vectors are generated and scored when using
HybridSearchVectorStorewith Qdrant and: view code..I am observing a significant mismatch between the query and the documents returned by the sparse vector search.
For example, after ingesting documents containing content about ASP.NET Core, a completely unrelated query such as:
"¿Qué es Colombia?"
can return chunks related to ASP.NET Core from the sparse vector store.
I would like to understand how this can happen at the vector/scoring level.
Specifically:
How does Qdrant/bm25 transform the document text into a sparse vector during ingestion?
How is the query transformed into a sparse vector?
Which terms and weights are stored in the sparse vector?
How is the similarity score between the query sparse vector and document sparse vector calculated?
Why can a document with apparently no relevant terms to the query obtain a significant non-zero score?
Does Ragbits apply any normalization or transformation to the BM25 score before returning the result?
Does HybridSearchVectorStore apply a score threshold to sparse results before combining them with dense results?
I am particularly interested in understanding the complete path:
Document
-> BM25 sparse vector
-> Qdrant indexed sparse vector
Query
-> BM25 sparse query vector
-> Qdrant sparse search
-> similarity score
-> HybridSearchVectorStore
The main concern is not that BM25 returns a non-zero score, but understanding why documents that appear completely unrelated to the query can obtain sufficiently high scores to be returned as relevant results.
Could you clarify whether this behavior is expected from the Qdrant/bm25 implementation, or whether Ragbits performs any additional processing that could explain this behavior?
Reference:
https://qdrant.tech/articles/sparse-vectors/
How can we reproduce it?
Relevant log output