Skip to content

v1.5.0 leaves ~60 GB of RAM unused and loses 13 points of expert hit rate: 18% slower decode than v1.4.0 on a 128 GB single-GPU host #885

Description

v1.5.0 is 17–19% slower at decode than v1.4.0 on this box, on both an int4-g64 and an E8/IQ3 container. Same machine, same hour, same frozen usage history, two replicates per arm on each version.

The interesting part is not the throughput, it is the resident set. v1.4.0 uses 106 GB of the 128 GB available. v1.5.0 uses 47 GB and never comes back for the rest. Expert hit rate falls with it, 97.4% → 84.4%. The engine is declining memory it has.

Setup

RTX 5080 / sm_120 / 16 GB · Core Ultra 9 285K · 128 GB DDR5-4400 · SK Hynix PCIe 4.0 NVMe · Windows · GLM-5.2 744B.

Both versions built from source with the same toolchain (nvcc -arch=sm_120, make colibri.exe CUDA_DLL=1 ARCH=native) — not the release zip, which carries no sm_120 coli_cuda.dll.

COLI_CUDA=1 CUDA_DENSE=1 CUDA_EXPERT_GB=4 CUDA_RESERVE_GB=1
PIPE=1 DIRECT=1
CACHE_ROUTE=1 ROUTE_J=3 ROUTE_M=16 ROUTE_ALPHA=0.5
USAGE_SAVE=0
--ngen 64 --temp 0, same prompt throughout

Preconditions recorded in both logs, because my first attempt at this was invalid without them — other GPU tenants had clamped the expert tier and the whole comparison was worthless:

VRAM  16303 MiB total, 1739 MiB used, 14239 MiB free
RAM   127.6 GB total / 107.6 GB free

.coli_usage is restored from a master copy before every single run and USAGE_SAVE=0, so the history is identical for all fourteen runs and no run can mutate it mid-flight. Warm-up discarded on each arm. v1.5.0 ran first, then the v1.4.0 binaries were reinstalled and the identical script re-run — same session, nothing else changed.

Result

arm v1.4.0 b085b48 v1.5.0 8f512fc delta
int4-g64 decode 1.21, 1.23 → 1.22 1.00, 1.01 → 1.005 −17.6%
E8/IQ3 decode 1.58, 1.58 → 1.58 1.27, 1.28 → 1.275 −19.3%
E8 speed profile J2/M24 1.66 1.46 −12.0%
g64 prefill, 19 tok 17.2, 17.8 s 22.6, 22.1 s +27%
E8 prefill, 19 tok 10.1, 10.0 s 11.4, 11.4 s +13%

Replicate spread is ≤0.02 tok/s on every arm on both versions, so this is well outside noise.

Where it comes from: the resident set, not the codec

v1.4.0 v1.5.0
g64 hit rate 93.3 / 93.5% 84.3 / 84.9%
E8 hit rate 97.4% 84.3 / 84.4%
E8 speed-profile hit 99.1% 90.9%
g64 RSS 105.8 GB 64.1 GB
E8 RSS 106.5 GB 47.5 GB
E8 pin / lru split 72.9% / 24.5% 57.9% / 26.4%

E8 loses 59 GB of resident set. The pinned share drops ~15 points while the LRU share gains only ~2, so the difference is not moved between tiers — it is simply not allocated. Hit rate follows, and throughput follows hit rate.

The two containers are unchanged files, the CUDA expert tier is the same size in both versions (156/1352 for g64, 228/1352 for E8, 3.3 GB), and the card is quiet. Nothing but the engine version differs.

Suspect

#815, perf(cache): preserve adaptive LRU during autopin (@bherald), which adds:

/* Automatic history pinning and the adaptive LRU share the expert RAM budget.
 * Preserve the LRU capacity affordable before pinning, up to the requested ... */
static double autopin_preserve_lru(double planned_pin, double expert_available,
                                   double lru_reserve){
    if(planned_pin<=0.0 || expert_available<=lru_reserve) return 0.0;
    double max_pin=expert_available-lru_reserve;

The intent is right — autopin was starving the LRU. But on this host the reserve appears to be computed far larger than what the LRU then uses, so max_pin is cut hard and the reserved remainder is never claimed by either tier. Both tiers end up small and ~60 GB sits idle. That matches the measured shape exactly: pin down 15 points, LRU up 2, total down 13.

I have not read enough of autopin_lru_reserve() to say which term is wrong, and I would rather hand over the measurement than a diagnosis I have not earned.

My .coli_usage carries 6,592,176 selections, which is large. If the reserve scales with history size this would show up more on well-used profiles than on fresh ones. That is a guess, and it is testable: the same sweep on a fresh .coli_usage would separate it.

What would settle it

  1. Print the computed lru_reserve and the resulting max_pin next to the existing [PIN] placement line. If the reserve reads ~60 GB on a 128 GB host, this is confirmed at a glance and by every user who hits it.
  2. The RSS figure already appears in the run summary — the regression is visible there. Anyone with both versions reproduces it in two runs.

Happy to re-run any arm, sweep PIN_GB, test a candidate patch, or run with a fresh usage history. Raw logs for all fourteen runs available.

Note for anyone reading this before it is fixed: v1.5.0 also carries eight security advisories, two of them in the model loader. This is a performance regression, not a reason to stay on v1.4.0 if you load models you did not build yourself.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugDifetto verificato nel codiceperformanceVelocità / tok-s / ottimizzazioni

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions