Comprehensive benchmarking tools and RAG examples for the SAGE framework
SAGE Benchmark provides a comprehensive suite of benchmarking tools and RAG (Retrieval-Augmented Generation) examples for evaluating SAGE framework performance. This package enables researchers and developers to:
- Benchmark RAG pipelines with multiple retrieval strategies (dense, sparse, hybrid)
- Compare vector databases (SageVDB, Milvus, FAISS) for RAG applications
- Evaluate multimodal retrieval with text, image, and video data
- Run reproducible experiments with standardized configurations and metrics
This package is designed for both research experiments and production system evaluation.
- Multiple RAG Implementations: Dense, sparse, hybrid, and multimodal retrieval
- Vector Database Support: SageVDB, Milvus, FAISS integration
- Experiment Framework: Automated benchmarking with configurable experiments
- Evaluation Metrics: Comprehensive metrics for RAG performance
- Sample Data: Included test data for quick start
- Extensible Design: Easy to add new benchmarks and retrieval methods
sage-benchmark/
├── src/
│ └── sage/
│ └── benchmark/
│ ├── __init__.py
│ └── benchmark_rag/ # RAG benchmarking
│ ├── __init__.py
│ ├── implementations/ # RAG implementations
│ │ ├── pipelines/ # RAG pipeline scripts
│ │ │ ├── qa_dense_retrieval_milvus.py
│ │ │ ├── qa_sparse_retrieval_milvus.py
│ │ │ ├── qa_multimodal_fusion.py
│ │ │ └── ...
│ │ └── tools/ # Supporting tools
│ │ ├── build_sagevdb_index.py
│ │ ├── build_milvus_dense_index.py
│ │ └── loaders/
│ ├── evaluation/ # Experiment framework
│ │ ├── pipeline_experiment.py
│ │ ├── evaluate_results.py
│ │ └── config/
│ ├── config/ # RAG configurations
│ └── data/ # Test data
│ # Future benchmarks:
│ # ├── benchmark_agent/ # Agent benchmarking
│ # └── benchmark_anns/ # ANNS benchmarking
├── tests/
├── pyproject.toml
└── README.md
Clone the repository with submodules and set up development environment:
# 1. Clone repository
git clone --recurse-submodules https://github.com/intellistream/sage-benchmark.git
cd sage-benchmark
# Or if already cloned, initialize submodules
./quickstart.sh
# 2. Install package with development dependencies
pip install -e ".[dev]"
# 3. Install pre-commit hooks (IMPORTANT for contributors)
pre-commit installThe quickstart.sh script will automatically:
- Initialize all Git submodules (LibAMM, SAGE-DB-Bench, sageData)
- Check environment and dependencies
- Display submodule status
Why install pre-commit? Pre-commit hooks automatically check code quality (formatting, import sorting, linting) before each commit, preventing CI/CD failures.
For the Q1–Q8 runner contract, output schema, CI inputs, aggregation flow, and failure troubleshooting, see the Q1–Q8 single-node benchmark guide.
Run the end-to-end benchmark pipeline in one command:
sage-benchmark-oneclickThis command executes:
- Full system benchmark (
Q1..Q8) - Aggregate + merge results for HF
- Upload to
intellistream/sage-benchmark-results - Refresh
sage-docsleaderboard data
Useful options:
# Quick smoke path
sage-benchmark-oneclick --quick
# Validate configs only, skip upload for local checks
sage-benchmark-oneclick --dry-run --skip-upload
# Explicit docs path
sage-benchmark-oneclick --docs-root /path/to/sage-docs
# If your environment does not expose sage.benchmark.benchmark_sage,
# provide an explicit benchmark entry command
sage-benchmark-oneclick --benchmark-command "python -m sage.benchmark.benchmark_sage --all --quick"If you prefer manual setup:
# Clone repository
git clone https://github.com/intellistream/sage-benchmark.git
cd sage-benchmark
# Initialize submodules (direct level only, not recursive)
git submodule update --init
# Install package
pip install -e .Or with development dependencies:
pip install -e ".[dev]"This repository uses Git submodules for external components:
- benchmark_amm (
src/sage/benchmark/benchmark_amm/) → LibAMM - benchmark_anns (
src/sage/benchmark/benchmark_anns/) → SAGE-DB-Bench - sage.data (
src/sage/data/) → sageData
All submodules track the main-dev branch and must be initialized before use.
The benchmark_rag module provides comprehensive RAG benchmarking capabilities:
Various RAG approaches for performance comparison:
Vector Databases:
- SageVDB: Native SAGE vector database backend
- Milvus: Dense, sparse, and hybrid retrieval
- FAISS: Efficient similarity search
Retrieval Methods:
- Dense retrieval (embeddings-based)
- Sparse retrieval (BM25, sparse vectors)
- Hybrid retrieval (combining dense + sparse)
- Multimodal fusion (text + image + video)
First, prepare your vector index:
# Build Milvus dense index
python -m sage.benchmark.benchmark_rag.implementations.tools.build_milvus_dense_indexTest individual RAG pipelines:
# Dense retrieval with Milvus
python -m sage.benchmark.benchmark_rag.implementations.pipelines.qa_dense_retrieval_milvus
# Sparse retrieval
python -m sage.benchmark.benchmark_rag.implementations.pipelines.qa_sparse_retrieval_milvus
# Hybrid retrieval (dense + sparse)
python -m sage.benchmark.benchmark_rag.implementations.pipelines.qa_hybrid_retrieval_milvusExecute full benchmark suite:
# Run comprehensive benchmark
python -m sage.benchmark.benchmark_rag.evaluation.pipeline_experiment
# Evaluate and generate reports
python -m sage.benchmark.benchmark_rag.evaluation.evaluate_resultsResults are saved in benchmark_results/:
experiment_TIMESTAMP/- Individual experiment runsmetrics.json- Performance metricscomparison_report.md- Comparison report
from sage.benchmark.benchmark_rag.implementations.pipelines import (
qa_dense_retrieval_milvus,
)
from sage.benchmark.benchmark_rag.config import load_config
# Load configuration
config = load_config("config_dense_milvus.yaml")
# Run RAG pipeline
results = qa_dense_retrieval_milvus.run_pipeline(query="What is SAGE?", config=config)
# View results
print(f"Retrieved {len(results)} documents")
for doc in results:
print(f"- {doc.content[:100]}...")from sage.benchmark.benchmark_rag.evaluation import PipelineExperiment
# Define experiment configuration
experiment = PipelineExperiment(
name="custom_rag_benchmark",
pipelines=["dense", "sparse", "hybrid"],
queries=["query1.txt", "query2.txt"],
metrics=["precision", "recall", "latency"],
)
# Run experiment
results = experiment.run()
# Generate report
experiment.generate_report(results)Configuration files are located in sage/benchmark/benchmark_rag/config/:
config_dense_milvus.yaml- Dense retrieval configurationconfig_sparse_milvus.yaml- Sparse retrieval configurationconfig_hybrid_milvus.yaml- Hybrid retrieval configurationconfig_dense_milvus.yaml- Milvus dense retrieval configuration
Experiment configurations in sage/benchmark/benchmark_rag/evaluation/config/:
experiment_config.yaml- Benchmark experiment settings
Test data is included in the package:
-
Benchmark Data (
benchmark_rag/data/):queries.jsonl- Sample queries for testingqa_knowledge_base.*- Knowledge base in multiple formats (txt, md, pdf, docx)sample/- Additional sample documents for testingsample/- Additional sample documents
-
Benchmark Config (
benchmark_rag/config/):experiment_config.yaml- RAG benchmark configurations
pytest packages/sage-benchmark/# Format code
black packages/sage-benchmark/
# Lint code
ruff check packages/sage-benchmark/For detailed documentation on each component:
- See
src/sage/benchmark/rag/README.mdfor RAG examples - See
src/sage/benchmark/benchmark_rag/README.mdfor benchmark details
- benchmark_agent: Agent system performance benchmarking
- benchmark_anns: Approximate Nearest Neighbor Search benchmarking
- benchmark_llm: LLM inference performance benchmarking
This package follows the same contribution guidelines as the main SAGE project. See the main
repository's CONTRIBUTING.md.
This project is licensed under the MIT License - see the LICENSE file for details.
- isage: Consolidated SAGE core for
sage.foundation,sage.stream, andsage.runtime - isagellm: External inference / gateway engine used by LLM and embedding benchmarks
- Optional adapters:
isage-rag,isage-neuromem,isage-vdbwhen a benchmark needs a specific retrieval or storage backend
- Documentation: https://intellistream.github.io/sage-docs/guides/packages/sage-benchmark/
- Issues: https://github.com/intellistream/SAGE/issues
- Discussions: https://github.com/intellistream/SAGE/discussions
Part of the SAGE Framework | Main Repository