- GPU: Hopper (H100/H20), Blackwell (B200), or newer — SM90+
- CUDA: 12.8+ toolkit installed (Blackwell/B200 requires 12.9+, see below)
- Python: 3.11+
- OS: Ubuntu 22.04 (tested)
- System packages: the NUMA dev headers (
numactl-develon RHEL/TencentOS,libnuma-devon Debian/Ubuntu) — required for thecore_engineAOT build (#include <numa.h>).install_deps.shinstalls these automatically; without them the installation fails withnuma.h: No such file or directory. - Model-specific (Kimi-K3 / Kimi-linear):
fla-core>=0.5.0(flash-linear-attention) provides theflaKimi-Delta-Attention kernels. It is listed inrequirements.txt, sopip install .(andinstall_deps.sh) pull it automatically — no separate step. - GitHub access: the repository is currently private — anonymous
git cloneand raw release-asset URLs fail (404). Authenticate first (e.g.gh auth login), clone viagh repo clone batchgen-project/batchgen, and fetch release wheels withgh release download(see Option C).
All three options below have been validated end-to-end on a fresh 2-node (16×H20) cluster
with Kimi-K2.5. Pick whichever fits your workflow. For Blackwell (B200), see the
Blackwell (B200) section for the additional BUILD_ARCH=sm100 knob.
cd BatchGen
docker build -t batchgen -f docker/Dockerfile .This builds everything in order: PyTorch → flash-attn 3 → FlashMLA → DeepGEMM → batchgen_kernels → batchgen. Takes ~40-60 min.
# 1. Create conda env
conda create -n batchgen python=3.11 -y
conda activate batchgen
# 2. Clone repo
git clone https://github.com/batchgen-project/batchgen.git
cd batchgen
# 3. Install everything
./scripts/install_deps.sh --allThis installs (in order):
- PyTorch 2.9.0+cu128 — from official wheels (~2 min)
- flash-attention 3 — built from source, Hopper only (~15-20 min)
- FlashMLA — built from source (~5-10 min)
- DeepGEMM — built from source (~5-10 min)
- batchgen_kernels — AOT-compiled CUDA extensions, 23 kernels (~7 min)
- batchgen — main package via
pip install .(~1 min)
Total: ~40-50 min on first install.
The production runtime is AOT-only: install_deps.sh builds and verifies
batchgen.core_engine during installation, and the server will not compile it
on first launch. See Runtime dependency contract for
the Torch/ CUDA ABI, worktree-isolation, tokenizer, FA3/FA2, UCX, and
fail-closed startup rules.
Pre-built wheels for Hopper GPUs (CUDA 12.8, PyTorch 2.9, Python 3.11) are
available on the GitHub Releases page.
Use a release that contains the complete validated wheel set: FlashAttention 3,
FlashMLA, SGL DeepGEMM, Apache TVM FFI, batchgen_kernels, and BatchGen. While
the repository is private, plain pip install <URL> returns 404; download the
wheels with an authenticated client first:
RELEASE_TAG="vX.Y.Z" # replace with the release tag you want to install
gh release download "$RELEASE_TAG" \
-R batchgen-project/batchgen \
-p '*.whl' \
-D ./wheels
pip install ./wheels/*.whlThe complete set must contain distributions named flash-attn-3, flash-mla,
sgl-deep-gemm==0.1.5.post3, apache-tvm-ffi==0.1.11,
batchgen-kernels, and batchgen. The installer validates these identities
and falls back to source for a missing or incompatible component.
To let the installer download and validate that set automatically:
conda create -n batchgen python=3.11 -y
conda activate batchgen
pip install torch==2.9.0+cu128 --index-url https://download.pytorch.org/whl/cu128
./scripts/install_deps.sh --all --release-tag "$RELEASE_TAG"When the complete wheel set is present, no source compilation is needed. Pip installs the remaining pure-Python dependencies (transformers, fastapi, etc.) from PyPI with the BatchGen wheel. The automatic installer builds only a missing or incompatible native component from source.
If you need to build wheels for a different environment:
bash scripts/build_wheels.sh --output-dir /path/to/wheelsThen upload to a GitHub Release:
gh release create v1.0.4.post1 --title "BatchGen v1.0.4.post1"
gh release upload v1.0.4.post1 /path/to/wheels/*.whlBatchGen runs on NVIDIA Blackwell (B200, compute capability 10.0). The engine
auto-detects the architecture (detect_gpu_arch() returns blackwell) and the
MLA models (DeepSeek-R1 / GLM-5 / Kimi-K2.5) route attention through the
FlashAttention-3 + FlashMLA path, the same as Hopper.
- CUDA toolkit /
nvcc≥ 12.9 — required to build FlashMLA's SM100 kernels. The Docker image (Option A) uses a CUDA 12.9 base for this reason. - PyTorch stays
2.9.0+cu128; only the build-timenvccneeds to be 12.9+.
Set BUILD_ARCH=sm100 so the Blackwell build paths are selected. This both
enables FlashMLA's SM100 kernels and builds batchgen_kernels for sm_100:
# Bare metal / conda (Option B):
BUILD_ARCH=sm100 ./scripts/install_deps.sh --all
# Docker (Option A) — pass it as a build arg:
docker build --build-arg BUILD_ARCH=sm100 -t batchgen:b200 -f docker/Dockerfile .With the default BUILD_ARCH=sm90a, the scripts/Dockerfile behave exactly as
the Hopper instructions above (FlashMLA SM100 disabled).
Use the B200 engine config, which sets "gpu_arch": "blackwell":
cd /root # not the source dir (see below)
python -m batchgen.launch_http_server \
--engine-config configurations/DeepSeek-R1/engine_config_B200_8.json \
...Note (kernel coverage): The dedicated SM100 port of the
batchgen_kernelsWGMMA MoE/attention kernels is tracked as a separate effort. Until it lands,BUILD_ARCH=sm100builds the SM80-class fused ops retargeted tosm_100; the WGMMA-only kernels are not yet built for Blackwell, and code paths that probe for them (e.g.is_qkv_wgmma_available()) fall back automatically. The MLA attention path (FlashAttention-3 / FlashMLA) is the intended Blackwell route.
After installing with pip install . (non-editable), you must not launch the
server from the BatchGen source directory. Python will find the source batchgen/
and batchgen_kernels/ directories before the installed site-packages, causing
import errors (e.g., missing compiled CUDA extensions).
# WRONG — source dir shadows installed packages
cd /path/to/BatchGen
python -m batchgen.launch_http_server ...
# CORRECT — run from any other directory
cd /root # or /tmp, or ~, etc.
python -m batchgen.launch_http_server ...This does not apply to Docker (Option A), where the source is the install target.
| Item | Detail |
|---|---|
| Install mode | pip install . (non-editable) — required for ray/production |
| Do not run from source dir | Source tree shadows installed packages (see above) |
| batchgen_kernels | Must use --no-build-isolation (needs installed PyTorch headers) |
| H20 GPUs | Set TORCH_CUDA_ARCH_LIST=9.0a before building kernels |
| Core engine | JIT-compiled at first server launch via ninja (automatic, ~5s) |
| No JIT for compute kernels | All 23 CUDA extensions are AOT-compiled in batchgen_kernels |
This discovers and imports every compiled _C_* extension shipped in the
installed batchgen_kernels package, so it stays correct as kernels are added:
python -c "
import glob, os, importlib
import batchgen_kernels
pkg_dir = os.path.dirname(batchgen_kernels.__file__)
exts = sorted(
f'batchgen_kernels.{os.path.basename(os.path.dirname(p))}.'
+ os.path.basename(p).split('.')[0]
for p in glob.glob(os.path.join(pkg_dir, '*', '_C_*.so'))
)
assert exts, 'No compiled batchgen_kernels extensions found'
for ext in exts:
importlib.import_module(ext)
print(f' {ext}: OK')
print(f'{len(exts)} batchgen_kernels extensions verified.')
import flash_attn_interface, flash_mla, deep_gemm
from importlib.metadata import version
assert version("sgl-deep-gemm") == "0.1.5.post3"
assert version("apache-tvm-ffi") == "0.1.11"
print('All dependencies verified.')
"PyTorch 2.9.0+cu128
├── flash-attention 3 (--no-build-isolation)
├── FlashMLA (--no-build-isolation)
├── SGL DeepGEMM (sgl-deep-gemm 0.1.5.post3 + apache-tvm-ffi 0.1.11)
├── batchgen_kernels (--no-build-isolation, 23 CUDAExtensions)
│ ├── SM90a: WGMMA kernels (MoE, QKV, routing)
│ └── SM80+: fused ops (RMSNorm, RoPE, dequant, MGN)
└── batchgen (pip install .)
└── core_engine (JIT at first launch, automatic)
See
docs/troubleshooting.mdfor the full list of commonly-met problems (build, runtime, and model bring-up).
core_engine is JIT-compiled by ninja at first launch and needs the NUMA dev headers.
Install them (sudo dnf install -y numactl-devel on RHEL/TencentOS, or
sudo apt-get install -y libnuma-dev on Debian/Ubuntu). install_deps.sh now does this
automatically. The runtime numactl/numactl-libs packages are not enough — the
-devel/-dev package ships numa.h.
flash-attention's setup.py tries to download a pre-built wheel matching your
torch version before building from source. No pre-built wheel exists for torch 2.9,
so the download hangs or times out. The install scripts already handle this by setting
FLASH_ATTENTION_FORCE_BUILD=TRUE, which skips the wheel download and builds from
source directly. If building manually:
cd flash-attention/hopper
FLASH_ATTENTION_FORCE_BUILD=TRUE pip install . --no-build-isolationFlashMLA's SM100 (Blackwell) codepath requires NVCC 12.9+. If you are building for
Hopper (the default, BUILD_ARCH=sm90a) the install scripts disable SM100 automatically:
FLASH_MLA_DISABLE_SM100=1 pip install . --no-build-isolationFor Blackwell, build with BUILD_ARCH=sm100 (which leaves SM100 enabled) and a CUDA
12.9+ toolkit. See the Blackwell (B200) section.
Ensure TORCH_CUDA_ARCH_LIST is set correctly (e.g., 9.0a for H20).
The batchgen_kernels build requires CUTLASS headers, which are vendored in the
repository at batchgen_kernels/3rd/cutlass/ (not a git submodule). If they are
missing, you likely have a shallow or partial checkout — re-clone the full repository.
If GitHub is unreliable, use Option C (pre-built wheels) which only requires
downloading .whl files. Alternatively, clone repos on a machine with better
connectivity and copy them over.