Skip to content

bug: How are BM25 sparse vectors generated and scored in HybridSearchVectorStore? #984

Description

@frankyjquintero

What happened?

Description

I am trying to understand how the sparse vectors are generated and scored when using HybridSearchVectorStore with Qdrant and: view code..

I am observing a significant mismatch between the query and the documents returned by the sparse vector search.

For example, after ingesting documents containing content about ASP.NET Core, a completely unrelated query such as:

"¿Qué es Colombia?"

can return chunks related to ASP.NET Core from the sparse vector store.

I would like to understand how this can happen at the vector/scoring level.

Specifically:

How does Qdrant/bm25 transform the document text into a sparse vector during ingestion?
How is the query transformed into a sparse vector?
Which terms and weights are stored in the sparse vector?
How is the similarity score between the query sparse vector and document sparse vector calculated?
Why can a document with apparently no relevant terms to the query obtain a significant non-zero score?
Does Ragbits apply any normalization or transformation to the BM25 score before returning the result?
Does HybridSearchVectorStore apply a score threshold to sparse results before combining them with dense results?

I am particularly interested in understanding the complete path:

Document
-> BM25 sparse vector
-> Qdrant indexed sparse vector

Query
-> BM25 sparse query vector
-> Qdrant sparse search
-> similarity score
-> HybridSearchVectorStore

The main concern is not that BM25 returns a non-zero score, but understanding why documents that appear completely unrelated to the query can obtain sufficiently high scores to be returned as relevant results.

Could you clarify whether this behavior is expected from the Qdrant/bm25 implementation, or whether Ragbits performs any additional processing that could explain this behavior?

Reference:
https://qdrant.tech/articles/sparse-vectors/

How can we reproduce it?

from qdrant_client import AsyncQdrantClient
from ragbits.core.vector_stores.qdrant import QdrantVectorStore
from ragbits.core.vector_stores.base import VectorStoreOptions
from ragbits.document_search import DocumentSearch, DocumentSearchOptions
from ragbits.core.embeddings.dense import LiteLLMEmbedder
from ragbits.core.embeddings.dense.fastembed import FastEmbedEmbedder
from ragbits.core.embeddings.sparse.fastembed import FastEmbedSparseEmbedder
from ragbits.core.vector_stores.hybrid import HybridSearchVectorStore
from ragbits.core.llms import LiteLLM
from ragbits.document_search import DocumentSearch
from ragbits.document_search.ingestion.parsers import DocumentParserRouter
from ragbits.document_search.documents.document import DocumentType
from ragbits.document_search.retrieval.rephrasers import LLMQueryRephraser
from ragbits.document_search.ingestion.parsers.docling import DoclingDocumentParser
from ragbits.core.embeddings.dense.litellm import LiteLLMEmbedderOptions
from ragbits_rag.config import config
from docling_core.transforms.chunker.hybrid_chunker import HybridChunker


def get_llm():
    return LiteLLM(
        model_name=f"huggingface/{config.llm_model}",
        api_base=config.llm_base_url,
        api_key=config.openai_api_key,
        use_structured_output=True,
    )

async def get_vector_store():
    store_prefix = "ragbits-rag"
    qdrant_client = AsyncQdrantClient(config.qdrant_host)

    dense_embedder = LiteLLMEmbedder(
        model_name=f"openai/{config.embedding_model}", 
        base_url=config.embedding_base_url,
        default_options=LiteLLMEmbedderOptions(encoding_format="float")
    )
    sparse_embedder = FastEmbedSparseEmbedder(
        model_name="Qdrant/bm25",
        use_gpu=False,
    )
    sparse_store = QdrantVectorStore(
        client=qdrant_client,
        embedder=sparse_embedder,
        index_name=f"{store_prefix}-sparse",
    )
    dense_store = QdrantVectorStore(client=qdrant_client, embedder=dense_embedder, index_name=store_prefix + "-dense")
    return HybridSearchVectorStore(
        dense_store,
        sparse_store,
    )
    

async def get_document_search():
    llm = get_llm()
    vector_store = await get_vector_store()
    chunker = HybridChunker(
        max_tokens=256,
        merge_peers=True,
    )

    # Configure document parsers
    parser_router = DocumentParserRouter({
        DocumentType.PDF: DoclingDocumentParser(chunker=chunker),
        DocumentType.DOCX: DoclingDocumentParser(chunker=chunker),
        DocumentType.HTML: DoclingDocumentParser(chunker=chunker),
        DocumentType.TXT: DoclingDocumentParser(chunker=chunker),
    })

    document_search = DocumentSearch(
        vector_store=vector_store,
        query_rephraser=LLMQueryRephraser(llm=llm),
        parser_router=parser_router,
        enricher_router=None
    )

    return document_search


uv run ragbits document-search --factory-path ragbits_rag.components:get_document_search ingest "web://https://www.webcolegios.com/file/51faaa.pdf"
uv run ragbits document-search --factory-path ragbits_rag.components:get_document_search ingest "web://https://pdfobject.com/pdf/sample.pdf"


uv run ragbits document-search --factory-path ragbits_rag.components:get_document_search search "que es colombia?"





import argparse
import asyncio

from ragbits_rag.components import get_document_search
from ragbits.core.vector_stores.base import VectorStoreOptions
from ragbits.document_search import DocumentSearchOptions


async def main(query: str, k: int, threshold: float | None):
    document_search = await get_document_search()

    options = DocumentSearchOptions(
        vector_store_options=VectorStoreOptions(
            k=k,
            score_threshold=threshold,
        )
    )

    results = await document_search.search(
        query,
        options=options,
    )
    print("#" * 75)
    for result in results:
        print("+" * 75)
        print(result)


if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("query")
    parser.add_argument("--k", type=int, default=10)
    parser.add_argument("--threshold", type=float)

    args = parser.parse_args()

    asyncio.run(main(
        args.query,
        args.k,
        args.threshold,
    ))


uv run python src/ragbits_rag/query_rag.py "que es colombia?" --k 10 --threshold 0.55

Relevant log output

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions