Skip to content

Benchmark: GLM-5.2 on Ryzen 7 9850X3D + RTX 5090, native Windows (0.89 tok/s median) #1091

Description

@Appie-NL

Machine datapoint

component detail
CPU AMD Ryzen 7 9850X3D, 8 cores / 16 threads
RAM 64 GB DDR5 4800MT/s
DISK 2TB WD_Black SN8100
GPU NVIDIA GeForce RTX 5090, 32607 MiB, driver 591.86
OS native Windows 11
model GLM-5.2 Colibri int4
Colibri commit 4ef9a9920dd4290b4ac26f74e73913fd18379bc9
build native Windows CUDA DLL, sm_120

Configuration

COLI_CUDA=1
CUDA_DENSE=1
CUDA_EXPERT_GB=auto
CUDA_RESERVE_GB=3
COLI_CUDA_PIPE=2
COLI_CUDA_ATTN=1
PIN=auto
PIN_GB=20
RAM_GB=52
CAP_RAISE=1
DRAFT=0
PIPE=1
PIPE_WORKERS=8
DIRECT=1
PILOT=1
PILOT_REAL=1
PILOT_WORKERS=4
cap=8
TEMP=0
NGEN=32

Fixed prompt: Continue this story: The lighthouse keeper climbed the stairs and saw something impossible in the fog.

The same fixed prompt was run three times. Existing .coli_usage history was retained, so these are warm, usage-trained runs rather than clean-install cold runs.

Decode results

run prefill (21 tokens) decode (32 tokens) throughput expert hit rate RSS
1 18.77 s 35.57 s 0.90 tok/s 38.1% 43.10 GB
2 18.93 s 36.15 s 0.89 tok/s 38.1% 43.10 GB
3 18.60 s 36.91 s 0.87 tok/s 38.2% 43.10 GB
median 18.77 s 36.15 s 0.89 tok/s 38.1% 43.10 GB

Additional runtime details:

  • CUDA resident set: 20.32 GB VRAM
  • CUDA expert tier: 478 resident experts / 10.15 GB
  • MTP compiled active, but DRAFT=0; measured 1.00-1.03 tokens/forward
  • First run reported 581.2 experts loaded per token

Disk

iobench out-00069.safetensors 19 64 8 <mode>:

mode throughput
buffered, 8 threads 4.80 GB/s
O_DIRECT, 8 threads 10.64 GB/s

Thread scaling with 128 reads: 1T 9.57, 2T 9.90, 4T 9.85, 8T 10.65, 12T 10.72, 16T 9.57 GB/s.

Runner compatibility note

The current c/tools/datapoint.py could not directly produce this GLM datapoint on native Windows. It invokes the older engine interface with positional cap and bits arguments, expects the older resident weights loaded in ... | RSS after load log line, and reports Windows RAM as unknown. The GLM engine instead takes the prompt through PROMPT, uses NGEN/MAX_NEW, and prints loaded in ... | resident dense: .... These measurements therefore use the documented fixed 32-token protocol directly against colibri.exe.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions