Skip to content

Benchmark: NVIDIA DGX Spark (GB10) 121GB — 11.1+ tok/s full-k8 CUDA (PLD); was 2.4 / 3.33 chat #161

Description

@VincentMarquez

Machine:** NVIDIA DGX Spark (NVIDIA_DGX_Spark) · Ubuntu 24.04 · Linux 6.17.0-1008-nvidia aarch64
CPU: 10× Cortex-X925 @ 3.9 GHz + 10× Cortex-A725 @ 2.8 GHz (20 cores)
Memory: 121 GB unified (CPU/GPU shared)
GPU: NVIDIA GB10 (sm_121), driver 580.126.09
Disk: NVMe (/dev/nvme0n1p2), model on local ext4
Model: GLM-5.2 colibri int4 · int8 MTP heads (sizes match matey-0 clone: 3527131672 / 5366238584 / 1065950496)
Engine base: JustVugg/colibri @ a78a06f + local experimental patches (CACHE_ROUTE / TRACE_EXPERTS / CUDA unified path)

Disk (README protocol)

c/iobench · 19 MB × 64 reads · 8 threads · shard out-00099.safetensors:

mode measured
buffered 4.25 GB/s
O_DIRECT 9.69 GB/s

How to read these numbers

Community table rows are usually stock routing (full top-8) with decode tok/s from the coli stats line after warm-up. This report has two tiers:

  1. Tier A — full top-8, no routing change (config num_experts_per_tok=8)
  2. Tier B — experimental CACHE_ROUTE (opt-in, not upstream default; ~14% expert substitution vs true top-8; cache-aware routing, arXiv:2412.00099 max-rank)

Also:

  • Decode tok/s = generation window only (coli decode tok/s). Overall includes prefill — do not use overall for ranking.
  • MTP=0 / DRAFT=0 for these speed cells.
  • Chat used TEMP=0.7. Output on the primes prompt was the correct first-20 list every turn.
  • Short context via :reset each turn (clears KV only; pin/LRU survive).

Tier A — full top-8, stock routing (TOPK unset, no CACHE_ROUTE)

Config: COLI_CUDA_UNIFIED=1 · CUDA_DENSE=1 · CUDA_GROUPED=1 · DIRECT=1 · PIPE=1 · large session LRU · learned AUTOPIN.

cell protocol decode tok/s hit notes
Timed ./glm PROFILE WARM=160, NGEN=48, greedy TEMP=0 2.39 82% best non-interactive full
Chat (warm session) primes prompt, :reset each turn ~2.08 ~82% chat ladder peak before CA

Example footer (pre–CACHE_ROUTE):

└─ 39 tok · 2.08 decode tok/s · … · hit 82% · RSS ~76 GB
   miss ~107/tok · I/O ~8.5s · expert ~3.2s · attention ~3.9s

───

Tier B — experimental CACHE_ROUTE (opt-in; not stock default)

CACHE_ROUTE=1 ROUTE_J=2 ROUTE_M=12
# keep true top-2; fill remaining 6 slots preferring pin∪LRU experts ranked in top-12

Sandbox predicted hit 0.81→0.955 and ~14% substitution. Live chat matches substitution and exceeds the hit target.

Interactive chat ladder

Prompt each turn:

List the first 20 prime numbers, one per line.

Protocol: one long-lived session · :reset before every turn · same prompt · full top-8 K · CACHE_ROUTE on.

┌──────┬─────┬──────────────┬─────┬──────┬──────┬──────────┐
│ turn │ tok │ decode tok/s │ hit │ swap │  I/O │ miss/tok │
├──────┼─────┼──────────────┼─────┼──────┼──────┼──────────┤
│    1 │  41 │         3.01 │ 95% │  15% │ 2.8s │     32.6 │
├──────┼─────┼──────────────┼─────┼──────┼──────┼──────────┤
│    2 │  39 │         3.16 │ 96% │  15% │ 2.1s │     25.1 │
├──────┼─────┼──────────────┼─────┼──────┼──────┼──────────┤
│    4 │  39 │         3.24 │ 96% │  14% │ 1.8s │     21.2 │
├──────┼─────┼──────────────┼─────┼──────┼──────┼──────────┤
│    6 │  39 │         3.29 │ 97% │  14% │ 1.6s │     18.4 │
├──────┼─────┼──────────────┼─────┼──────┼──────┼──────────┤
│ best │  41 │         3.33 │ 97% │  14% │ 1.5s │     18.0 │
└──────┴─────┴──────────────┴─────┴──────┴──────┴──────────┘

Best (Tier B): 3.33 decode tok/s · hit 97% · swap 14% · RSS 78.5 GB · wall ~18s for 41 tokens.

Verbatim footer:

└─ 41 tok · 3.33 decode tok/s · 2.27 overall · hit 97% · swap 14% · RSS 78.5 GB · 18s
   miss 18.0/tok (0.34 GB/tok) · MTP 1.00 tok/fw (0% accept) · I/O 1.5s · expert 3.2s · attention 4.1s
   CUDA 112.5% of expert calls (27675 calls) · 1.00 rows/expert
   cache admission 80 accepted · 660 scan-pollution rejects

Output quality on this prompt: correct first-20 primes every turn. Not a full ./coli bench quality gate — still pending before defaulting CACHE_ROUTE.

Timed one-shot (heavy warm, less favorable than multi-turn chat LRU):

48 tokens in 18.00s (2.67 tok/s decode) | expert hit rate 93.5% | swap 15.0%
PROFILE: expert-disk 3.852s | expert-matmul 3.833s | attention 6.155s | other 4.162s

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarkDatapoint di misurazione hardwarecudaBackend CUDA/NVIDIA

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions