Machine datapoint
| component |
detail |
| CPU |
AMD Ryzen 7 9850X3D, 8 cores / 16 threads |
| RAM |
64 GB DDR5 4800MT/s |
| DISK |
2TB WD_Black SN8100 |
| GPU |
NVIDIA GeForce RTX 5090, 32607 MiB, driver 591.86 |
| OS |
native Windows 11 |
| model |
GLM-5.2 Colibri int4 |
| Colibri commit |
4ef9a9920dd4290b4ac26f74e73913fd18379bc9 |
| build |
native Windows CUDA DLL, sm_120 |
Configuration
COLI_CUDA=1
CUDA_DENSE=1
CUDA_EXPERT_GB=auto
CUDA_RESERVE_GB=3
COLI_CUDA_PIPE=2
COLI_CUDA_ATTN=1
PIN=auto
PIN_GB=20
RAM_GB=52
CAP_RAISE=1
DRAFT=0
PIPE=1
PIPE_WORKERS=8
DIRECT=1
PILOT=1
PILOT_REAL=1
PILOT_WORKERS=4
cap=8
TEMP=0
NGEN=32
Fixed prompt: Continue this story: The lighthouse keeper climbed the stairs and saw something impossible in the fog.
The same fixed prompt was run three times. Existing .coli_usage history was retained, so these are warm, usage-trained runs rather than clean-install cold runs.
Decode results
| run |
prefill (21 tokens) |
decode (32 tokens) |
throughput |
expert hit rate |
RSS |
| 1 |
18.77 s |
35.57 s |
0.90 tok/s |
38.1% |
43.10 GB |
| 2 |
18.93 s |
36.15 s |
0.89 tok/s |
38.1% |
43.10 GB |
| 3 |
18.60 s |
36.91 s |
0.87 tok/s |
38.2% |
43.10 GB |
| median |
18.77 s |
36.15 s |
0.89 tok/s |
38.1% |
43.10 GB |
Additional runtime details:
- CUDA resident set: 20.32 GB VRAM
- CUDA expert tier: 478 resident experts / 10.15 GB
- MTP compiled active, but
DRAFT=0; measured 1.00-1.03 tokens/forward
- First run reported 581.2 experts loaded per token
Disk
iobench out-00069.safetensors 19 64 8 <mode>:
| mode |
throughput |
| buffered, 8 threads |
4.80 GB/s |
| O_DIRECT, 8 threads |
10.64 GB/s |
Thread scaling with 128 reads: 1T 9.57, 2T 9.90, 4T 9.85, 8T 10.65, 12T 10.72, 16T 9.57 GB/s.
Runner compatibility note
The current c/tools/datapoint.py could not directly produce this GLM datapoint on native Windows. It invokes the older engine interface with positional cap and bits arguments, expects the older resident weights loaded in ... | RSS after load log line, and reports Windows RAM as unknown. The GLM engine instead takes the prompt through PROMPT, uses NGEN/MAX_NEW, and prints loaded in ... | resident dense: .... These measurements therefore use the documented fixed 32-token protocol directly against colibri.exe.
Machine datapoint
4ef9a9920dd4290b4ac26f74e73913fd18379bc9Configuration
Fixed prompt:
Continue this story: The lighthouse keeper climbed the stairs and saw something impossible in the fog.The same fixed prompt was run three times. Existing
.coli_usagehistory was retained, so these are warm, usage-trained runs rather than clean-install cold runs.Decode results
Additional runtime details:
DRAFT=0; measured 1.00-1.03 tokens/forwardDisk
iobench out-00069.safetensors 19 64 8 <mode>:Thread scaling with 128 reads: 1T 9.57, 2T 9.90, 4T 9.85, 8T 10.65, 12T 10.72, 16T 9.57 GB/s.
Runner compatibility note
The current
c/tools/datapoint.pycould not directly produce this GLM datapoint on native Windows. It invokes the older engine interface with positionalcapandbitsarguments, expects the olderresident weights loaded in ... | RSS after loadlog line, and reports Windows RAM as unknown. The GLM engine instead takes the prompt throughPROMPT, usesNGEN/MAX_NEW, and printsloaded in ... | resident dense: .... These measurements therefore use the documented fixed 32-token protocol directly againstcolibri.exe.