v1.5.0 is 17–19% slower at decode than v1.4.0 on this box, on both an int4-g64 and an E8/IQ3 container. Same machine, same hour, same frozen usage history, two replicates per arm on each version.
The interesting part is not the throughput, it is the resident set. v1.4.0 uses 106 GB of the 128 GB available. v1.5.0 uses 47 GB and never comes back for the rest. Expert hit rate falls with it, 97.4% → 84.4%. The engine is declining memory it has.
Setup
RTX 5080 / sm_120 / 16 GB · Core Ultra 9 285K · 128 GB DDR5-4400 · SK Hynix PCIe 4.0 NVMe · Windows · GLM-5.2 744B.
Both versions built from source with the same toolchain (nvcc -arch=sm_120, make colibri.exe CUDA_DLL=1 ARCH=native) — not the release zip, which carries no sm_120 coli_cuda.dll.
COLI_CUDA=1 CUDA_DENSE=1 CUDA_EXPERT_GB=4 CUDA_RESERVE_GB=1
PIPE=1 DIRECT=1
CACHE_ROUTE=1 ROUTE_J=3 ROUTE_M=16 ROUTE_ALPHA=0.5
USAGE_SAVE=0
--ngen 64 --temp 0, same prompt throughout
Preconditions recorded in both logs, because my first attempt at this was invalid without them — other GPU tenants had clamped the expert tier and the whole comparison was worthless:
VRAM 16303 MiB total, 1739 MiB used, 14239 MiB free
RAM 127.6 GB total / 107.6 GB free
.coli_usage is restored from a master copy before every single run and USAGE_SAVE=0, so the history is identical for all fourteen runs and no run can mutate it mid-flight. Warm-up discarded on each arm. v1.5.0 ran first, then the v1.4.0 binaries were reinstalled and the identical script re-run — same session, nothing else changed.
Result
| arm |
v1.4.0 b085b48 |
v1.5.0 8f512fc |
delta |
| int4-g64 decode |
1.21, 1.23 → 1.22 |
1.00, 1.01 → 1.005 |
−17.6% |
| E8/IQ3 decode |
1.58, 1.58 → 1.58 |
1.27, 1.28 → 1.275 |
−19.3% |
| E8 speed profile J2/M24 |
1.66 |
1.46 |
−12.0% |
| g64 prefill, 19 tok |
17.2, 17.8 s |
22.6, 22.1 s |
+27% |
| E8 prefill, 19 tok |
10.1, 10.0 s |
11.4, 11.4 s |
+13% |
Replicate spread is ≤0.02 tok/s on every arm on both versions, so this is well outside noise.
Where it comes from: the resident set, not the codec
|
v1.4.0 |
v1.5.0 |
| g64 hit rate |
93.3 / 93.5% |
84.3 / 84.9% |
| E8 hit rate |
97.4% |
84.3 / 84.4% |
| E8 speed-profile hit |
99.1% |
90.9% |
| g64 RSS |
105.8 GB |
64.1 GB |
| E8 RSS |
106.5 GB |
47.5 GB |
| E8 pin / lru split |
72.9% / 24.5% |
57.9% / 26.4% |
E8 loses 59 GB of resident set. The pinned share drops ~15 points while the LRU share gains only ~2, so the difference is not moved between tiers — it is simply not allocated. Hit rate follows, and throughput follows hit rate.
The two containers are unchanged files, the CUDA expert tier is the same size in both versions (156/1352 for g64, 228/1352 for E8, 3.3 GB), and the card is quiet. Nothing but the engine version differs.
Suspect
#815, perf(cache): preserve adaptive LRU during autopin (@bherald), which adds:
/* Automatic history pinning and the adaptive LRU share the expert RAM budget.
* Preserve the LRU capacity affordable before pinning, up to the requested ... */
static double autopin_preserve_lru(double planned_pin, double expert_available,
double lru_reserve){
if(planned_pin<=0.0 || expert_available<=lru_reserve) return 0.0;
double max_pin=expert_available-lru_reserve;
The intent is right — autopin was starving the LRU. But on this host the reserve appears to be computed far larger than what the LRU then uses, so max_pin is cut hard and the reserved remainder is never claimed by either tier. Both tiers end up small and ~60 GB sits idle. That matches the measured shape exactly: pin down 15 points, LRU up 2, total down 13.
I have not read enough of autopin_lru_reserve() to say which term is wrong, and I would rather hand over the measurement than a diagnosis I have not earned.
My .coli_usage carries 6,592,176 selections, which is large. If the reserve scales with history size this would show up more on well-used profiles than on fresh ones. That is a guess, and it is testable: the same sweep on a fresh .coli_usage would separate it.
What would settle it
- Print the computed
lru_reserve and the resulting max_pin next to the existing [PIN] placement line. If the reserve reads ~60 GB on a 128 GB host, this is confirmed at a glance and by every user who hits it.
- The
RSS figure already appears in the run summary — the regression is visible there. Anyone with both versions reproduces it in two runs.
Happy to re-run any arm, sweep PIN_GB, test a candidate patch, or run with a fresh usage history. Raw logs for all fourteen runs available.
Note for anyone reading this before it is fixed: v1.5.0 also carries eight security advisories, two of them in the model loader. This is a performance regression, not a reason to stay on v1.4.0 if you load models you did not build yourself.
v1.5.0 is 17–19% slower at decode than v1.4.0 on this box, on both an int4-g64 and an E8/IQ3 container. Same machine, same hour, same frozen usage history, two replicates per arm on each version.
The interesting part is not the throughput, it is the resident set. v1.4.0 uses 106 GB of the 128 GB available. v1.5.0 uses 47 GB and never comes back for the rest. Expert hit rate falls with it, 97.4% → 84.4%. The engine is declining memory it has.
Setup
RTX 5080 / sm_120 / 16 GB · Core Ultra 9 285K · 128 GB DDR5-4400 · SK Hynix PCIe 4.0 NVMe · Windows · GLM-5.2 744B.
Both versions built from source with the same toolchain (nvcc
-arch=sm_120,make colibri.exe CUDA_DLL=1 ARCH=native) — not the release zip, which carries no sm_120coli_cuda.dll.Preconditions recorded in both logs, because my first attempt at this was invalid without them — other GPU tenants had clamped the expert tier and the whole comparison was worthless:
.coli_usageis restored from a master copy before every single run andUSAGE_SAVE=0, so the history is identical for all fourteen runs and no run can mutate it mid-flight. Warm-up discarded on each arm. v1.5.0 ran first, then the v1.4.0 binaries were reinstalled and the identical script re-run — same session, nothing else changed.Result
b085b488f512fcReplicate spread is ≤0.02 tok/s on every arm on both versions, so this is well outside noise.
Where it comes from: the resident set, not the codec
E8 loses 59 GB of resident set. The pinned share drops ~15 points while the LRU share gains only ~2, so the difference is not moved between tiers — it is simply not allocated. Hit rate follows, and throughput follows hit rate.
The two containers are unchanged files, the CUDA expert tier is the same size in both versions (
156/1352for g64,228/1352for E8, 3.3 GB), and the card is quiet. Nothing but the engine version differs.Suspect
#815,
perf(cache): preserve adaptive LRU during autopin(@bherald), which adds:The intent is right — autopin was starving the LRU. But on this host the reserve appears to be computed far larger than what the LRU then uses, so
max_pinis cut hard and the reserved remainder is never claimed by either tier. Both tiers end up small and ~60 GB sits idle. That matches the measured shape exactly: pin down 15 points, LRU up 2, total down 13.I have not read enough of
autopin_lru_reserve()to say which term is wrong, and I would rather hand over the measurement than a diagnosis I have not earned.My
.coli_usagecarries 6,592,176 selections, which is large. If the reserve scales with history size this would show up more on well-used profiles than on fresh ones. That is a guess, and it is testable: the same sweep on a fresh.coli_usagewould separate it.What would settle it
lru_reserveand the resultingmax_pinnext to the existing[PIN] placementline. If the reserve reads ~60 GB on a 128 GB host, this is confirmed at a glance and by every user who hits it.RSSfigure already appears in the run summary — the regression is visible there. Anyone with both versions reproduces it in two runs.Happy to re-run any arm, sweep
PIN_GB, test a candidate patch, or run with a fresh usage history. Raw logs for all fourteen runs available.Note for anyone reading this before it is fixed: v1.5.0 also carries eight security advisories, two of them in the model loader. This is a performance regression, not a reason to stay on v1.4.0 if you load models you did not build yourself.