Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
265 commits
Select commit Hold shift + click to select a range
b5a8bac
refactor(plan): isolate unqualified GPU discovery
Kenneth-Javier Aug 10, 2026
80370ed
feat(plan): discover Windows AMD GPUs via hipInfo
Kenneth-Javier Aug 10, 2026
ad8a75b
fix(doctor): recognize the expected Windows GPU backend
Kenneth-Javier Aug 10, 2026
9ffa28b
fix(cli): exclude unqualified GPUs from auto-enable
Kenneth-Javier Aug 10, 2026
dd6c22d
docs(plan): clarify Windows unified-memory discovery
Kenneth-Javier Aug 10, 2026
36378ea
fix(doctor): keep the Windows linkage probe portable in tests
Kenneth-Javier Aug 11, 2026
2e65c13
feat(v4): dual-SSD mirror (COLI_MODEL_MIRROR) for the expert store
dcutugno Aug 10, 2026
725ead5
fix(v4/mirror): use linear-index hash for balanced expert routing
dcutugno Aug 12, 2026
bb92d80
fix(v4/mirror): count primary direct reads in mirror stats
dcutugno Aug 12, 2026
22ef387
fix(v4/mirror): update test for v4_read_expert_record with rep parameter
dcutugno Aug 12, 2026
3cc2420
DeepSeek V4: the CUDA kernels, as a tier rather than an engine
ZacharyZcR Aug 6, 2026
763dac7
Remove vendored DeepSeek V4 CUDA paths
ZacharyZcR Aug 12, 2026
ddf5cbb
metal: widen the fmt-gate notice to all fused-bound tensors, share th…
monotophic Aug 13, 2026
dc99ce8
metal: mode-aware kv_b remedy, accurate MTP rationale, notice-test ha…
monotophic Aug 13, 2026
574aac1
docs: recompute FORMATS.md line anchors for post-#989-wave dev
monotophic Aug 14, 2026
d81284a
test(v4): end-to-end OpenAI wire tests for V4 DSML tool calling
AcciaiCS Aug 14, 2026
042d21f
performance: use coli_v4_route_bf16 in moe_token
inqode-lars Aug 14, 2026
3d96376
fix(v4/mirror): review follow-ups — serve telemetry, per-replica O_DI…
dcutugno Aug 14, 2026
6aa3069
Merge remote-tracking branch 'origin/dev' into pr/v4-ssd-mirror
dcutugno Aug 14, 2026
e53018c
fix(v4): pack hot pinned experts outside state->mutex (issue #900)
gouravkargwal Aug 14, 2026
7065f20
Merge pull request #947 from gouravkargwal/ci/metal-test
JustVugg Aug 14, 2026
0f335e5
Merge pull request #958 from Blakeolson21/fix/openmp-team-launch-paths
JustVugg Aug 14, 2026
a35e30a
fix(docker): package the gateway's DSML module
Blakeolson21 Aug 14, 2026
345bc6c
fix(convert): publish resumable shards atomically
Blakeolson21 Aug 14, 2026
7072c04
fix(convert): publish OLMoE shards atomically
Blakeolson21 Aug 14, 2026
efce7a9
Merge origin/dev into fix/v4-usage-route-trace
terrizoaguimor Aug 15, 2026
d1653ca
fix(mirror): recognize Kimi K3 expert shards
terrizoaguimor Aug 15, 2026
e27c37e
Merge pull request #969 from terrizoaguimor/fix/v4-usage-route-trace
JustVugg Aug 15, 2026
2487342
docs(api): map tool calling support by engine
terrizoaguimor Aug 15, 2026
600a3e0
build: add sm_121 gencode to the portable fatbin and pin -ftz=false
monotophic Aug 15, 2026
268c74f
cuda: warp-per-row fmt=8 kernels with reference-mirroring accumulation
monotophic Aug 15, 2026
9aa32f2
tests: fmt=8 warp-kernel oracle (sweep, census/tail shapes, S-invaria…
monotophic Aug 15, 2026
a6fb69a
bench: fmt=8 old-vs-warp kernel bandwidth harness (JSON, fp8-bench ta…
monotophic Aug 15, 2026
71f921b
docs: register COLI_CUDA_F8_WARP in the environment table
monotophic Aug 15, 2026
3e9af13
tests: keep the old-vs-new census compare informational
monotophic Aug 15, 2026
9bf565d
bench: query roofline via device attributes (cudaDeviceProp members g…
monotophic Aug 15, 2026
87d4d98
refactor: extract shared routed-expert FFN helper
lineape Aug 15, 2026
cb73d1a
cuda: dense fmt=8 honors COLI_CUDA_F8_WARP; unify vec/byte product se…
monotophic Aug 15, 2026
1b6de08
tests: grouped accumulation-convention bite and vec/byte load-path pa…
monotophic Aug 15, 2026
1419dff
bench: roofline from BENCH_PEAK_GBPS only; report raw memory attributes
monotophic Aug 15, 2026
7783377
build: gate compute_121 on nvcc >= 12.9; run fp8_warp_test before mxfp4
monotophic Aug 15, 2026
669b488
tools: commit run_f8_bench.sh (BENCH_PEAK_GBPS=273 default for GB10)
monotophic Aug 15, 2026
8ab17dc
docs: COLI_CUDA_F8_WARP entry matches per-vendor default and mode sem…
monotophic Aug 15, 2026
aabe371
tests: dense accumulation-convention bite through coli_cuda_matmul
monotophic Aug 15, 2026
7aaf9d2
build: close the compute_121 gate when nvcc version is unparsable
monotophic Aug 15, 2026
17bfc3e
tests: drop direct cuda_runtime.h includes (hipcc has no such header)
monotophic Aug 15, 2026
84c85e5
tests: qualify host isnan as std::isnan (HIP cmath host overload wart)
monotophic Aug 15, 2026
b6af434
fix(quant): pin the grouped-int4 group accumulate as an fma
8PotatoChip8 Aug 15, 2026
a53e108
fix(cuda): mxfp4 group scale applied per 32-group when the scale is inf
monotophic Aug 15, 2026
84c6154
test(cuda): make the mxfp4 compare NaN-aware
monotophic Aug 15, 2026
d35fe81
test(cuda): ISA-robust exp-255 shapes and mixed-scale coverage
monotophic Aug 15, 2026
fa8b211
fix(convert): declare numpy in the OLMoE guard, and correct the OLMoE…
sami7969 Aug 16, 2026
cc2d039
fix(build): detect Windows from uname when $(OS) is unset (MSYS2/MinGW)
sami7969 Aug 16, 2026
dc3f8fd
fix(windows): set _WIN32_WINNT floor for FILE_ID_INFO under MinGW
sami7969 Aug 16, 2026
7da954a
Merge origin/dev into fix/metal-i4-gpu-path
aaristov Aug 16, 2026
84beb3e
coli: fix stale colibri.c line reference after the dev merge
aaristov Aug 16, 2026
e230ae6
docs: refresh the Discord invite (the old one expired)
JustVugg Aug 16, 2026
6d5ceff
Merge pull request #1033 from terrizoaguimor/docs/tool-calling-matrix
JustVugg Aug 16, 2026
f1bb9c3
Merge pull request #1028 from terrizoaguimor/fix/k3-mirror-plan
JustVugg Aug 16, 2026
c243b4c
Merge pull request #1027 from Blakeolson21/blake/atomic-olmoe-shards
JustVugg Aug 16, 2026
a49509a
Merge pull request #1026 from Blakeolson21/blake/atomic-fp8-shards
JustVugg Aug 16, 2026
1b1a822
Merge pull request #1025 from Blakeolson21/blake/docker-v4-dsml
JustVugg Aug 16, 2026
255e8aa
Merge pull request #1051 from JustVugg/fix/discord-invite
JustVugg Aug 16, 2026
7a671f0
Merge remote-tracking branch 'origin/dev' into pr/v4-ssd-mirror
dcutugno Aug 16, 2026
ff30465
Merge dev into the DeepSeek V4 CUDA tier (resolve c/Makefile)
JustVugg Aug 16, 2026
a720c80
Merge pull request #1041 from monotophic/fix/mxfp4-exp255-refmirror
JustVugg Aug 16, 2026
4f86057
Merge pull request #1037 from monotophic/cuda-f8-warp
JustVugg Aug 16, 2026
abbb8de
Merge pull request #931 from Kenneth-Javier/fix/amd-discovery-windows
JustVugg Aug 16, 2026
db0720c
Merge dev (keep $(NVCC_STD) and -ftz=false)
JustVugg Aug 16, 2026
e36a1c7
Merge pull request #1023 from gouravkargwal/fix/v4-rows16-pack-outsid…
JustVugg Aug 16, 2026
b1d3403
Merge pull request #1054 from JustVugg/zach/deepseek-v4-cuda-tier
JustVugg Aug 16, 2026
9371b17
Merge pull request #988 from dcutugno/pr/v4-ssd-mirror
JustVugg Aug 16, 2026
c64fc66
Merge pull request #1010 from gcaponi/v4-tools-e2e-tests
JustVugg Aug 16, 2026
1a8e091
Merge pull request #1038 from 8PotatoChip8/fix/fma-contract-grouped-int4
JustVugg Aug 16, 2026
a568db1
Merge pull request #830 from dpanelli/swa-kv-ring-buffer
JustVugg Aug 16, 2026
3d5200d
Merge pull request #934 from Nanetnounou/pr/neon-e8
JustVugg Aug 16, 2026
03e8677
Merge pull request #827 from monotophic/kvb/fmt-gate-notice-r3
JustVugg Aug 16, 2026
af48fe8
olmoe: make IDOT opt-in on x86 and NEON, and add the test that was mi…
Aug 16, 2026
51638e8
feat(v4): prefix checkpoints (memory + disk), resumable 4096-token pr…
dcutugno Aug 16, 2026
62be01b
feat(v4/serve): prefix hint in SUBMIT, hold admission until the engin…
dcutugno Aug 16, 2026
623f6c1
docs(v4): prefill segments/chunks/checkpoints, env additions, what CP…
dcutugno Aug 16, 2026
d44b096
fix(plan): match DeepSeek V4 expert tensor naming (layers.N.ffn.exper…
dcutugno Aug 16, 2026
35349e1
fix(windows): make coli stop actually find the serve, and stop orphan…
JustVugg Aug 16, 2026
4eddc62
feat(v4/cuda): the DeepSeek V4 CUDA kernels the engine tier needs — g…
dcutugno Aug 16, 2026
66fca1e
feat(v4): the CUDA tier wired into the engine — every stage CPU-canon…
dcutugno Aug 16, 2026
5eb9b32
docs(v4): CUDA tier — build (Windows DLLs, Linux CUDA=1), serve line,…
dcutugno Aug 16, 2026
ca657ce
test: fix the Windows job test — never put the runner in the job
JustVugg Aug 16, 2026
dd39c4e
Merge pull request #1055 from dcutugno/pr/v4-cuda-tier
JustVugg Aug 16, 2026
9e10ebd
Merge pull request #1059 from JustVugg/fix/windows-stop-and-orphan-en…
JustVugg Aug 16, 2026
2410c6e
Merge pull request #1061 from nadiadatepe-eng/fix/olmoe-idot-default-off
JustVugg Aug 16, 2026
85951ef
Merge pull request #1047 from sami7969/fix/windows-build-blockers
JustVugg Aug 16, 2026
0578bba
Merge pull request #1035 from lineape/pr551-prefactor-shared-ffn
JustVugg Aug 16, 2026
fe3879c
refactor(core): make model families registry-owned
terrizoaguimor Aug 16, 2026
8c5f046
fix(doctor): accept engine-supported I64/F8 dtypes in --deep validator
jerome-benoit Aug 16, 2026
0fd5228
flake: build every supported engine, unpin COLI_ENGINE
jerome-benoit Aug 16, 2026
12368b7
Merge pull request #1063 from terrizoaguimor/refactor/family-registry
JustVugg Aug 16, 2026
eeaa6b6
Merge pull request #1065 from jerome-benoit/feat/flake-all-engines
JustVugg Aug 17, 2026
4d51576
Merge pull request #1064 from jerome-benoit/fix/doctor-deep-dtypes
JustVugg Aug 17, 2026
36e8397
Merge pull request #1048 from sami7969/fix/olmoe-numpy-and-size
JustVugg Aug 17, 2026
ed75a25
Merge pull request #965 from bherald/feat/k3-mmap-prepared-weights
JustVugg Aug 17, 2026
b84154d
fix(test): stub numpy in the OLMoE converter import mock — dev CI is red
JustVugg Aug 17, 2026
1cc0be2
Merge pull request #1067 from JustVugg/fix/olmoe-converter-test-numpy…
JustVugg Aug 17, 2026
0998991
feat: add distributed expert workers
gauravsaini Jul 23, 2026
4626a50
refactor: reuse openai_server HTTP hardening in cluster registry
lineape Aug 15, 2026
a380a04
feat: worker init parity — st_init_multi + COLI_MODEL_DIRS, clean unr…
lineape Aug 15, 2026
cf872b0
fix: route cluster worker through expert_ffn and rotate fmt=6 input
lineape Aug 15, 2026
fa92f96
fix: invert cluster_connect_one success check in cluster_init
lineape Aug 15, 2026
e270be1
test: COLIEX01 cluster wire-protocol round-trip test
lineape Aug 15, 2026
43ed2a7
fix: reload worker expert slot when slab is missing (use-after-free a…
lineape Aug 15, 2026
edcfd05
tools: quantize glm_tiny routed experts to fmt=6 (E8/IQ3) and fmt=4 (…
lineape Aug 15, 2026
99fec32
tools: de-duplicate fmt6/fmt4 quantizer in make_glm_oracle.py
lineape Aug 15, 2026
4982f66
test: token-exact parity gate — local vs cluster-delegated sharding (#7)
lineape Aug 15, 2026
cea3779
ci: exercise the cluster parity gate against the fmt6/fmt4 fixtures
lineape Aug 15, 2026
6608dca
docs: correct the fmt=6 down-rotation accounting in the expert_ffn co…
lineape Aug 15, 2026
723c1bb
feat: allow cross-host workers to register via --allowed-host
lineape Aug 15, 2026
7875c21
fix: check the model directory before resolving its family
lineape Aug 17, 2026
6bfbd3e
build: install cluster.py; complete test_cluster_protocol prerequisites
lineape Aug 17, 2026
2ecb5d4
refactor(core): finish registry-owned family dispatch
terrizoaguimor Aug 17, 2026
89365ae
v4: pluggable expert-store backend registry
8PotatoChip8 Aug 11, 2026
ca17810
perf: byte-identical CPU/startup optimizations
zh-Processor Aug 17, 2026
58e6c44
Merge pull request #1068 from terrizoaguimor/refactor/family-registry…
JustVugg Aug 17, 2026
f4dcbed
Merge pull request #1017 from weber-software/dev
JustVugg Aug 17, 2026
d4894f0
Merge pull request #1036 from lineape/pr551-rebased
JustVugg Aug 17, 2026
5e8e99d
perf: hoist expert activation quantization to layer level (bit-identi…
JustVugg Aug 17, 2026
804732e
docs(k3): clarify mmap RSS accounting
bherald Aug 17, 2026
c3e3edd
fix(server): account admitted client cancellations
bherald Aug 17, 2026
e5ff623
fix(k3): cancel prefill between layers
bherald Aug 17, 2026
8512368
Merge pull request #964 from 8PotatoChip8/v4/expert-store-registry
JustVugg Aug 17, 2026
c945cd5
Merge pull request #1071 from JustVugg/perf/qrow-hoist
JustVugg Aug 17, 2026
a7f2f64
perf(kimi_k3): hoist expert activation quantization to layer level (b…
JustVugg Aug 17, 2026
5d0986c
Merge pull request #1075 from JustVugg/perf/qrow-hoist-k3
JustVugg Aug 18, 2026
853a90a
perf(deepseek_v4): thread-local scratch for qdq activations (bit-iden…
JustVugg Aug 18, 2026
19fba39
Merge branch 'dev' into main
JustVugg Aug 18, 2026
45e0f99
fix: extend the qdq scratch to the DUAL and BATCH units too
JustVugg Aug 18, 2026
7c6cdc8
perf(deepseek_v4): qdq the attention input once — wq_a and wkv share it
JustVugg Aug 18, 2026
f70b1ed
Merge pull request #1076 from JustVugg/perf/qrow-hoist-v4
JustVugg Aug 18, 2026
377a6eb
Merge pull request #1077 from JustVugg/perf/qrow-hoist-v4b
JustVugg Aug 18, 2026
57eeb73
Merge pull request #1058 from zh-Processor/main
JustVugg Aug 18, 2026
17be735
wip(k1): planar int4 kernels — dot_i4p_u, matmul_i4p_idot, matmul_i4p…
JustVugg Aug 18, 2026
63b8036
test(serve): freeze OLMoE wire framing
terrizoaguimor Aug 18, 2026
d8ea56f
perf(colibri): K1 — plane-nibble int4 layout + unsigned-VNNI dot (bit…
JustVugg Aug 18, 2026
73cc7e1
Merge pull request #1072 from bherald/docs/k3-mmap-rss-accounting
JustVugg Aug 18, 2026
aa05364
Merge pull request #1074 from bherald/fix/server-cancel-accounting
JustVugg Aug 18, 2026
6132a67
Merge pull request #1073 from bherald/fix/k3-prefill-cancellation
JustVugg Aug 18, 2026
01d59ae
fix(inkling): IDOT opt-in — the fast path is x86-only and not token-e…
JustVugg Aug 18, 2026
fae9553
Merge pull request #1079 from JustVugg/perf/k1-planar
JustVugg Aug 18, 2026
d2d9fd6
feat(moe): FUSED3=1 opt-in AVX2 expert matmul — vectorized activation…
outtodata Aug 14, 2026
a2ce84e
ci: ARM oracle job + integer-kernel bit-exactness gate (closes #1081)
JustVugg Aug 18, 2026
62b8d63
Merge pull request #1080 from JustVugg/fix/inkling-idot-optin
JustVugg Aug 18, 2026
438d09a
Merge pull request #1082 from JustVugg/perf/fused3-olmoe
JustVugg Aug 18, 2026
6c63091
fix(web): auto-connect in useEffect, atlas from engine origin, drop i…
JustVugg Aug 18, 2026
a8e0a5f
docs: CHANGELOG catch-up v1.2.0–v1.6.2; SUMMARY.md was an experiment …
JustVugg Aug 18, 2026
2d2d92d
Merge pull request #1083 from JustVugg/ci/arm-oracle
JustVugg Aug 18, 2026
feb9dba
fix(k1): planar off on non-AVX2 builds; drop WIP loop; f32 planar in …
JustVugg Aug 18, 2026
6104bbb
refactor(serve): share framing with OLMoE
terrizoaguimor Aug 18, 2026
b04d602
Merge pull request #1084 from JustVugg/fix/web-triplet
JustVugg Aug 18, 2026
1de2939
Merge pull request #1085 from JustVugg/docs/changelog-summary
JustVugg Aug 18, 2026
0ad0ca7
Merge pull request #1086 from JustVugg/fix/k1-followup
JustVugg Aug 18, 2026
3cb832f
perf(colibri): K2 — 1x4 union tile in the planar IDOT matmul (bit-ide…
JustVugg Aug 18, 2026
19afc84
fix(test): the f32 planar assertion must mirror planar_on() exactly
JustVugg Aug 18, 2026
4da41bf
Merge pull request #1089 from JustVugg/fix/exactness-avx512-gate
JustVugg Aug 18, 2026
256bbe9
Merge pull request #1088 from JustVugg/perf/k2-union-tile
JustVugg Aug 18, 2026
a0d7cfb
Merge pull request #1087 from terrizoaguimor/refactor/serve-codec-olmoe
JustVugg Aug 18, 2026
6a33455
test(kimi): freeze serve wire framing
terrizoaguimor Aug 18, 2026
467bec4
build(v4/cuda): fetch the pinned DeepGEMM sm120 headers on first buil…
dcutugno Aug 16, 2026
fc52d73
feat(qwen36): add Qwen3.6-35B-A3B engine (CPU): hybrid Gated Attentio…
kreuzzelg Jul 30, 2026
9c901b5
feat(qwen36): make engine drivable by coli + SERVE=1 gateway protocol
minne100 Aug 1, 2026
5a0715e
feat(qwen36): group-scaled int4 (gs64) containers — converter flag + …
kreuzzelg Aug 1, 2026
24d3c89
docs(qwen36): user-facing page — prebuilt containers as the easy path
kreuzzelg Aug 3, 2026
2ff73e0
fix(qwen36): serve mode must not require an argv prompt file
kreuzzelg Aug 3, 2026
342f104
fix(qwen36): serve EOS ids from the tokenizer, not the 151k-vocab con…
kreuzzelg Aug 3, 2026
6a9f293
feat(qwen36): qwen chat-template family in the gateway
kreuzzelg Aug 3, 2026
c30c545
fix(qwen36): heap per-thread attention score rows — 8k context smashe…
kreuzzelg Aug 13, 2026
3252675
fix(qwen36): expert_get must not evict an in-flight slot (A5)
kreuzzelg Aug 16, 2026
9ae0c8a
ci: build and ship qwen36, and gate it against a tiny oracle (B5 + B5b)
kreuzzelg Aug 17, 2026
e84005e
ci(qwen36-tiny-check): make the gate deterministic and actually exact
kreuzzelg Aug 17, 2026
4e0dc88
fix(qwen36): validate every config dimension before the forward pass …
kreuzzelg Aug 17, 2026
ca96e12
fix(qwen36): refuse a container whose tensors do not match the config…
kreuzzelg Aug 17, 2026
cb8524b
Merge pull request #1056 from dcutugno/pr/v4-deepgemm
JustVugg Aug 18, 2026
cbc72fe
refactor(serve): share framing with Kimi K3
terrizoaguimor Aug 18, 2026
e755280
qwen36: the remaining cheap items from the review's C list
kreuzzelg Aug 18, 2026
67c7490
qwen36: drop the packed int4 buffers this engine never reads
kreuzzelg Aug 18, 2026
7e226fc
docs(qwen36): fix two stale references
kreuzzelg Aug 18, 2026
046842e
ci(qwen36-tiny-check): the sanitizer step asserts memory safety, not …
kreuzzelg Aug 18, 2026
921ce46
perf(colibri): K3-lite — parallel silu and the down-side half of the …
JustVugg Aug 18, 2026
6784aa8
Merge pull request #712 from kreuzzelg/qwen36-engine
JustVugg Aug 18, 2026
34deee3
feat(qwen36): CUDA VRAM expert tier — heat-based placement across GPU…
kreuzzelg Jul 30, 2026
0d0d1ed
Merge pull request #1093 from JustVugg/perf/k3-lite
JustVugg Aug 18, 2026
8c0aa8c
Merge pull request #1090 from terrizoaguimor/refactor/serve-codec-kimi
JustVugg Aug 18, 2026
7d1d19e
Merge pull request #713 from kreuzzelg/qwen36-cuda-tier
JustVugg Aug 18, 2026
88a4a82
perf(colibri): K1b — grouped planar IDOT for gs64 containers, opt-in …
JustVugg Aug 18, 2026
ad79236
Merge pull request #1094 from JustVugg/perf/k1b-gs-idot
JustVugg Aug 18, 2026
5fa92da
test(v4): freeze serve wire framing
terrizoaguimor Aug 18, 2026
1545fac
feat(planner): OLMoE PlannerGeometry adapter (#1066)
benmaster82 Aug 18, 2026
8cbebce
refactor(serve): share framing with DeepSeek V4
terrizoaguimor Aug 18, 2026
2b7a892
fix(cuda): barrier the absorb softmax reduction before shared red[] i…
monotophic Aug 17, 2026
07e48a7
test(cuda): cover run-to-run reproducibility of the batch and ragged …
monotophic Aug 17, 2026
80fbea4
fix(c): check #798's alloc/snprintf results on the checkpoint-load path
monotophic Aug 18, 2026
ebc68da
perf(deepseek_v4): expert-loader pool default 3 -> 9 lanes (measured …
JustVugg Aug 19, 2026
31debb7
Merge pull request #1097 from JustVugg/perf/v4-lanes-default
JustVugg Aug 19, 2026
93226cf
fix(loader): refuse duplicate tensor names across indexed shards
monotophic Aug 18, 2026
48e4574
fix(arm64): probe dotprod by compiling, prefer armv8.2-a base (#1104)
SebaWag Aug 19, 2026
10af5e9
fix(cuda): add the __syncwarp() grouped_s4_wmma's siblings carry (clo…
JustVugg Aug 19, 2026
93271fd
Merge pull request #1109 from SebaWag/fix/arm64-dotprod-probe
JustVugg Aug 19, 2026
2b5ae0e
Merge pull request #1111 from JustVugg/fix/cuda-syncwarp
JustVugg Aug 19, 2026
d0f32d1
chore: version 1.7.0
JustVugg Aug 19, 2026
4459b61
Metal Kimi K3: aligned KDA state/window bufs, wrap-once buffer cache,…
RDouglasSharp Aug 2, 2026
a851e5b
docs: Kimi K3 Metal tuning (M5 Max) + ignore build junk
RDouglasSharp Aug 2, 2026
f130c4e
Merge pull request #1098 from monotophic/fix/cuda-absorb-syncthreads
JustVugg Aug 19, 2026
3bad5a2
fix(cuda): stop fmt=8 tensor_free from overcounting freed scale bytes
monotophic Aug 18, 2026
5823580
fix(cuda): stop coli_cuda_tensor_bytes from overcounting fmt=8/fmt=6 …
monotophic Aug 18, 2026
ee5610b
Merge pull request #1113 from JustVugg/metal-k3-rebase
JustVugg Aug 19, 2026
9480d88
test(cuda): wire test_fp8_cuda into cuda-test and add independent foo…
monotophic Aug 18, 2026
f08bb08
Merge branch 'dev' into fix/metal-i4-gpu-path
JustVugg Aug 19, 2026
c09fb52
Merge pull request #1095 from benmaster82/feat/olmoe-planner-1066
JustVugg Aug 19, 2026
78ea979
Merge pull request #1096 from terrizoaguimor/refactor/serve-codec-v4
JustVugg Aug 19, 2026
c54b28d
Merge pull request #1101 from monotophic/fix/798-loader-guards
JustVugg Aug 19, 2026
f2f8253
Merge pull request #1106 from monotophic/fix/duplicate-tensor-name-re…
JustVugg Aug 19, 2026
bc46e96
test(inkling): freeze serve wire framing
terrizoaguimor Aug 19, 2026
c379f47
refactor(serve): share framing with Inkling
terrizoaguimor Aug 19, 2026
eae3d5c
merge dev into release branch
JustVugg Aug 19, 2026
d0ef133
docs: CHANGELOG — Metal Kimi K3, the correctness batch, planner and c…
JustVugg Aug 19, 2026
9b6e949
Merge pull request #1114 from JustVugg/fix/cuda-fmt8-rebased
JustVugg Aug 19, 2026
b00688b
fix(cache): make LRU victim selection respect a lowered ecap
ZacharyZcR Aug 19, 2026
269440d
fix(telemetry): honour USAGE_SAVE=0 in every engine
ZacharyZcR Aug 19, 2026
692e61e
Merge pull request #829 from aaristov/fix/metal-i4-gpu-path
JustVugg Aug 19, 2026
a4b3748
docs: CHANGELOG — GPU observability counters (#829)
JustVugg Aug 19, 2026
849f960
planner: add OLMoE geometry adapter (issue #1066)
SebaWag Aug 19, 2026
de99057
planner: add Kimi K3 hybrid geometry adapter (issue #1066)
SebaWag Aug 19, 2026
d7a4890
planner: add Inkling hybrid GQA/sconv geometry adapter (issue #1066)
SebaWag Aug 19, 2026
655dde7
planner: add DeepSeek V4 geometry adapter (issue #1066) -- completes …
SebaWag Aug 19, 2026
863bf1e
chore: apply the v1.6.2 warning-cleanup patches from #1032
ZacharyZcR Aug 19, 2026
7612433
fix(cuda): set weights_owned before the H2D copy so failed uploads fr…
monotophic Aug 19, 2026
01b70a1
test(cuda): keep the memcpy injection hook inside backend_gpu_compat.…
monotophic Aug 19, 2026
97bf97f
test(cuda): alias cudaMemcpyKind/cudaErrorInvalidValue per vendor for…
monotophic Aug 19, 2026
aabf942
Merge pull request #1122 from ZacharyZcR/fix/1039-usage-save
JustVugg Aug 19, 2026
0419d23
Merge pull request #1123 from ZacharyZcR/fix/1032-warnings
JustVugg Aug 19, 2026
e0e3cc4
Merge pull request #1115 from JustVugg/fix/cuda-weights-owned-rebased
JustVugg Aug 19, 2026
4a1f763
docs: CHANGELOG — the CUDA fix batch, USAGE_SAVE, warning cleanup
JustVugg Aug 19, 2026
909f2b4
Merge pull request #1121 from ZacharyZcR/fix/1034-rss-guard-cap
JustVugg Aug 19, 2026
f7d4754
docs: CHANGELOG — LRU victim selection vs lowered ecap (#1121)
JustVugg Aug 19, 2026
940ea50
Merge pull request #1112 from JustVugg/release/v1.7.0
JustVugg Aug 19, 2026
56b0303
test(family_registry): production planners report real geometry (#1066)
SebaWag Aug 19, 2026
072f832
fix(tools): support deepseek_v4 CLI dispatch and zero-write Linux cac…
SyedYousufFaizan Aug 19, 2026
070dbe8
ci: retry cancelled toolchain checks
terrizoaguimor Aug 19, 2026
06b7479
Merge pull request #1103 from SebaWag/feat/olmoe-planner-geometry
JustVugg Aug 19, 2026
9627aff
docs: CHANGELOG — planner geometry adapters for every family (#1103)
JustVugg Aug 19, 2026
b1806bd
Merge pull request #1116 from terrizoaguimor/refactor/serve-codec-ink…
JustVugg Aug 19, 2026
c5d7afa
docs: CHANGELOG — the serve codec now covers every engine (#1116)
JustVugg Aug 19, 2026
90dbbcb
Merge pull request #1125 from SyedYousufFaizan/fix/datapoint-v4-fadvise
JustVugg Aug 19, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
227 changes: 226 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,11 @@ jobs:
run: |
docker run --rm colibri:ci --version
docker run --rm colibri:ci --help > /dev/null
# The gateway has runtime modules that launcher-only commands do not
# import. Exercise that boundary so the advertised serve path cannot
# ship with a missing Python file.
docker run --rm --entrypoint python3 colibri:ci -c \
"import openai_server; print('gateway imports')"
- name: The engine binary is in the image and is executable
run: docker run --rm --entrypoint /app/colibri colibri:ci --help 2>&1 | head -5 || true
- name: Lint both Dockerfiles
Expand Down Expand Up @@ -185,14 +190,30 @@ jobs:
run: |
cd c
rc=0
ENGINES="colibri inkling kimi_k3 olmoe"
ENGINES="colibri inkling kimi_k3 olmoe qwen36"
if [ "${{ matrix.v4 }}" = "1" ]; then ENGINES="$ENGINES deepseek-v4"; fi
for t in $ENGINES; do
echo "::group::$t"
make $t || { echo "FAILED: $t"; rc=1; }
echo "::endgroup::"
done
exit $rc
- if: matrix.os == 'macos-latest'
name: Metal backend tests
# The Metal backend suite is its own target (metal-test) and is not
# part of `make check`: TEST_BINS is derived from tests/test_*.$(EXE)
# rules, and backend_metal_test is built/run by the metal-test target
# itself. Nothing in CI exercised it until issue #940 shipped a
# compile error in it, so run it explicitly on the Apple Silicon
# runners.
# The runner's Apple Paravirtual device hangs on Metal submissions
# (observed: 47 min vs a 5-6 min normal run), so the test harness
# detects it and skips GPU execution with an explicit line; this
# timeout is the backstop that turns any residual hang into a red X.
timeout-minutes: 10
run: |
cd c
make metal-test

# macOS is the newest COLI_V4_SUPPORTED platform and the only one whose
# engine coverage lives solely in this job (Linux has the v4-tiny job in
Expand Down Expand Up @@ -296,6 +317,111 @@ jobs:
-Xcompiler=-Wall,-Wextra
echo "inkling CUDA syntax check passed"

qwen36-tiny-check:
name: Qwen3.6 tiny oracle (token-exact + ASan/UBSan)
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
cache: pip
cache-dependency-path: c/tools/oracle-requirements.txt
- name: Install torch (CPU) + transformers
run: pip install -r c/tools/oracle-requirements.txt
- name: Tiny Qwen3.6-shaped fixture + full-hybrid oracle
run: |
cd c
# Same layout as the 35B model (10 x (3 x DeltaNet -> MoE, 1 x Attention
# -> MoE)) at toy dimensions, so the converter and the engine treat it
# exactly like the real one. --mode full exercises BOTH layer kinds;
# attention_only would leave the DeltaNet path untested.
# The reference comes from make_qwen36_tiny.py, not make_qwen36_oracle.py:
# the oracle encodes a text prompt through AutoTokenizer.from_pretrained(),
# and this fixture is synthetic -- weights, no tokenizer. The tiny script
# already has the model in memory and generates greedily from fixed ids.
# --ref-mode full leaves BOTH layer kinds active; the default
# attention_only replaces the 30 DeltaNet layers with identity, which is
# what Phase 1 computed and would leave three quarters of the engine
# untested.
python3 tools/make_qwen36_tiny.py --out qwen36_tiny --ref-mode full \
--emit-ref qwen36_tiny/ref_full.json
# ebits=8 keeps the expert quantization error below anything that could
# flip a greedy argmax here -- measured. The DENSE weights are the ones
# that matter: the engine quantizes them to int8 by default, the torch
# reference does not, and that alone cost 7 of 16 tokens. The runs below
# therefore set COLI_DENSE_I8=0, which is what the accuracy A/B in
# 07_Tests uses for the same reason. Without it this gate compares two
# different models and reddens on the luck of the draw.
python3 tools/convert_qwen36.py --model qwen36_tiny --out qwen36_tiny_c --ebits 8
- name: Token-exact against the oracle, at several cache capacities
run: |
cd c
make qwen36
# cap=1 evicts on every routed expert, which is where slot bookkeeping
# breaks; cap=8 never would. That is not hypothetical here: the
# in-flight eviction fallback this engine inherited was exactly such a
# bug, and it only shows when cap is smaller than the number of loads
# in flight. The engine exits non-zero on any token mismatch.
for cap in 1 2 8; do
echo "::group::cap=$cap"
COLI_DENSE_I8=0 SNAP=qwen36_tiny_c \
./qwen36 "$cap" 8 qwen36_tiny/ref_full.json
echo "::endgroup::"
done
- name: Same run under ASan + UBSan
run: |
cd c
# Token-exactness alone would not have caught the config-driven heap
# overflows this engine shipped with: they do not necessarily change
# the output. The sanitizers are the half of this gate that watches
# memory, and PILOT=1 puts concurrent expert loads against the cache
# so the slot paths are exercised, not just walked past.
#
# This step asserts MEMORY SAFETY, not token-exactness -- the three
# runs above already own that. -fsanitize changes inlining and
# vectorization, so this is a differently-compiled binary, and float
# reassociation can flip a token wherever the margin is thin. Demanding
# exactness from it conflates two goals and made this job red on a
# docs-only commit: the normal build matched 16/16, the sanitizer build
# did not. So: run it, ignore a token mismatch, fail on any sanitizer
# diagnostic.
make clean >/dev/null 2>&1 || true
make qwen36 EXTRA_CFLAGS="-fsanitize=address,undefined -fno-omit-frame-pointer -g"
ASAN_OPTIONS=detect_leaks=0 UBSAN_OPTIONS=print_stacktrace=1 \
COLI_DENSE_I8=0 SNAP=qwen36_tiny_c PILOT=1 WIDE=2 \
./qwen36 1 8 qwen36_tiny/ref_full.json > san.log 2>&1 || true
if grep -qE "ERROR: AddressSanitizer|runtime error:" san.log; then
echo "FAIL: sanitizer diagnostic under PILOT"; cat san.log; exit 1
fi
echo "sanitizers clean:"; grep -E "Matching tokens" san.log || true
- name: A malformed container is refused, not read
run: |
cd c
# The case ASan caught in review: config.json and qwen36_meta.json ship
# in the same container and disagreed on the layer count, so is_attn was
# sized from one and written from the other. Reproduced without the
# guard as a heap-buffer-overflow WRITE at load_meta; with it the engine
# refuses. A well-formed fixture cannot cover this, which is why it gets
# its own step -- and it asserts the REFUSAL, so a guard that silently
# stops guarding fails here.
cp -r qwen36_tiny_c qwen36_tiny_bad
python3 - <<'PY'
import json
p = "qwen36_tiny_bad/config.json"
cfg = json.load(open(p))
cfg["num_hidden_layers"] = 4 # meta says 8
json.dump(cfg, open(p, "w"))
PY
if COLI_DENSE_I8=0 SNAP=qwen36_tiny_bad ./qwen36 8 8 \
qwen36_tiny/ref_full.json > bad.log 2>&1; then
echo "FAIL: engine accepted a container whose two config files disagree"
cat bad.log; exit 1
fi
grep -q "config.json says 4 layers" bad.log || {
echo "FAIL: refused, but not for the reason under test"; cat bad.log; exit 1; }
echo "refused as expected:"; grep '^\[cfg\]' bad.log

inkling-oracle:
name: Inkling oracle (token-exact vs transformers)
runs-on: ubuntu-latest
Expand Down Expand Up @@ -360,6 +486,14 @@ jobs:
run: pip install -r c/tools/oracle-requirements.txt
- name: Build colibri
run: make -C c colibri
- name: Integer-kernel exactness vs pure-C reference (#1081)
# Integer kernels have no rounding excuse: bit-equality on every ISA.
# The same binary check runs on the ARM job below — this is the gate
# that would have caught the olmoe/qwen36/inkling IDOT class.
run: |
cd c
gcc -O3 -march=native -fopenmp -I. tests/test_int_kernel_exact.c -o /tmp/tike -lm
/tmp/tike
- name: Generate the glm_tiny fixture
run: cd c && python3 tools/make_glm_oracle.py
- name: Token-exact oracle (teacher forcing)
Expand All @@ -381,6 +515,42 @@ jobs:
tests.test_inefficiency.TinyEfficiencyTest.test_profile_phases_present_and_nonneg \
tests.test_inefficiency.TinyEfficiencyTest.test_disk_wait_not_dominant

# The token-exact parity gate from #7 (c/tests/test_cluster_sharding.py): the
# engine's local CPU run must already reproduce the transformers oracle
# token-exactly, and delegating the routed-expert shard to a cluster worker
# must not perturb a single position. It exercises the fmt=6 (E8/IQ3,
# rotation-bearing) path and its fmt=4 (grouped int4, no-rotation) control.
#
# Not in `make check`: the two fixtures are gitignored (regenerated, not
# committed -- see .gitignore) and need the pinned torch/transformers to
# build, the same rationale as inkling-oracle and efficiency above. The test
# skips cleanly when the fixtures are absent, so leaving fixture generation
# out of CI would turn the gate into a silent no-op on a clean host -- the
# worker path would never actually run. This job regenerates both fixtures
# first so the gate is exercised, not dormant.
#
# Socket safety: the test spawns a cluster worker subprocess and a
# coordinator, each binding a 127.0.0.1 port obtained from _free_port()
# (bind to port 0), so it is safe on a shared runner.
cluster-parity:
name: Cluster parity gate (fmt6/fmt4, token-exact)
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
cache: pip
cache-dependency-path: c/tools/oracle-requirements.txt
- name: Install torch (CPU) + transformers
run: pip install -r c/tools/oracle-requirements.txt
- name: Build colibri
run: make -C c colibri
- name: Generate the fmt6/fmt4 quantized fixtures
run: cd c && python3 tools/make_glm_oracle.py --fmt6 && python3 tools/make_glm_oracle.py --fmt4
- name: Cluster parity gate (token-exact)
run: cd c && python3 -m unittest -v tests.test_cluster_sharding

engine-hip-syntax:
name: HIP syntax check
runs-on: ubuntu-latest
Expand Down Expand Up @@ -470,7 +640,30 @@ jobs:
path-type: inherit
install: >-
make
git
mingw-w64-ucrt-x86_64-gcc
# The DeepSeek V4 DeepGEMM flavour compiles against a pinned checkout
# that tools/fetch_deepgemm.sh makes on first use (nothing vendored).
# CI cannot compile it (sm_120a needs CUDA >= 12.8; this job pins 12.6)
# but it does exercise the fetch — pin verification + MSVC patch — and
# caches the tree keyed on the pin so a bump is the only cache miss.
- name: DeepGEMM pin (for the cache key)
id: deepgemm-pin
shell: bash
run: echo "pin=$(sed -n 's/^DEEPGEMM_PIN ?= //p' c/Makefile | head -1)" >> "$GITHUB_OUTPUT"
- name: Cache the pinned DeepGEMM checkout
uses: actions/cache@v4
with:
path: c/third_party/deepgemm
key: deepgemm-${{ steps.deepgemm-pin.outputs.pin }}-${{ hashFiles('c/patches-deepgemm-sm120-msvc.patch') }}
- name: make deepgemm-fetch (pinned checkout + MSVC patch, no compile)
shell: msys2 {0}
run: |
cd c
make deepgemm-fetch
test -f third_party/deepgemm/deep_gemm/include/deep_gemm/impls/sm120_bf16_gemm.cuh || { echo "fetch_deepgemm produced no sm120 headers" >&2; exit 1; }
# second run must be a no-op (idempotency)
make deepgemm-fetch | grep -q "already at" || { echo "fetch_deepgemm is not idempotent" >&2; exit 1; }
- name: make cuda-dll (nvcc + MSVC host)
shell: msys2 {0}
run: |
Expand Down Expand Up @@ -537,3 +730,35 @@ jobs:
'tests.test_convert_positioned_write'
)
python -m unittest -v @suites

# ARM oracle (#1081): every tiny-oracle job above runs on x86, so NEON
# branches were never compiled, let alone executed — which is how the
# olmoe (#1044) / qwen36 (#712 review) / inkling (#1080) IDOT class shipped.
# This job builds every engine on arm64 (NEON compile coverage), runs the
# integer-kernel bit-exactness gate with the NEON branches live, and replays
# the glm_tiny teacher-forcing oracle against a fixture generated on THIS
# runner (same-machine torch reference, so no cross-ISA float excuses).
oracle-arm:
name: ARM (engines + NEON kernel exactness + tiny oracle)
runs-on: ubuntu-24.04-arm
timeout-minutes: 25
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
cache: pip
cache-dependency-path: c/tools/oracle-requirements.txt
- name: Install torch (CPU) + transformers
run: pip install -r c/tools/oracle-requirements.txt
- name: Build every engine (NEON branches must compile)
run: make -C c colibri inkling kimi_k3 olmoe
- name: Integer-kernel exactness, NEON branches live (#1081)
run: |
cd c
gcc -O3 -mcpu=native -fopenmp -I. tests/test_int_kernel_exact.c -o /tmp/tike -lm
/tmp/tike
- name: Generate the glm_tiny fixture on this runner
run: cd c && python3 tools/make_glm_oracle.py
- name: Token-exact oracle on ARM (teacher forcing)
run: cd c && SNAP=./glm_tiny TF=1 COLI_TEMP=0 ./colibri 64 16 16
34 changes: 23 additions & 11 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -62,10 +62,10 @@ jobs:
- name: Build engines
run: |
cd c
for t in colibri inkling kimi_k3 olmoe; do
for t in colibri inkling kimi_k3 olmoe qwen36; do
make $t ${{ matrix.make_args }}
done
ls -lh colibri${{ matrix.ext }} inkling${{ matrix.ext }} kimi_k3${{ matrix.ext }} olmoe${{ matrix.ext }}
ls -lh colibri${{ matrix.ext }} inkling${{ matrix.ext }} kimi_k3${{ matrix.ext }} olmoe${{ matrix.ext }} qwen36${{ matrix.ext }}

# DeepSeek V4 has its own Makefile and its own target name (`deepseek-v4`,
# producing `deepseek_v4`), so it never joined the loop above -- and #858 is
Expand Down Expand Up @@ -120,6 +120,10 @@ jobs:
cp c/inkling${{ matrix.ext }} dist/inkling${{ matrix.ext }}
cp c/kimi_k3${{ matrix.ext }} dist/kimi_k3${{ matrix.ext }}
cp c/olmoe${{ matrix.ext }} dist/olmoe${{ matrix.ext }}
# qwen36 -- plain name, like the others: engine_for() resolves "qwen" to
# a "qwen36" binary next to coli. Unconditional: the engine is portable C
# with no platform gate, so every archive can carry it.
cp c/qwen36${{ matrix.ext }} dist/qwen36${{ matrix.ext }}
# deepseek_v4 -- underscore, not the hyphen of the make target: engine_for()
# looks for "deepseek_v4" next to coli. Guarded, not unconditional, because
# the macos-arm64 archive has no such engine to ship (#858).
Expand All @@ -135,6 +139,7 @@ jobs:
cp c/version.py dist/
cp c/openai_server.py dist/
cp c/v4_dsml.py dist/ # openai_server imports v4_dsml (vendored DeepSeek V4 DSML primitives); without it the packaged server won't import
cp c/family_registry.py dist/ # authoritative model-family dispatch/capability registry
cp c/resource_plan.py dist/
cp c/doctor.py dist/
cp c/autotune.py dist/
Expand Down Expand Up @@ -167,7 +172,7 @@ jobs:
# a green build, green tests, and an archive with no engine for the model the
# user actually had. Presence is not enough -- assert they are executable, and
# that the launcher's own resolver finds them where it looks (next to coli).
SIBLINGS="inkling kimi_k3 olmoe"
SIBLINGS="inkling kimi_k3 olmoe qwen36"
# deepseek_v4 only where the engine builds (COLI_V4_SUPPORTED); asserting it
# unconditionally would fail the macos-arm64 archive for shipping an engine
# that platform cannot have.
Expand All @@ -184,11 +189,9 @@ jobs:
# fine. The check was wrong, not the artifact.
python3 -c "import ast; ast.parse(open('tools/k3_tokenizer.py', encoding='utf-8').read())" \
|| { echo "FAIL: packaged k3_tokenizer.py does not parse"; exit 1; }
# COLI_EXPECT rather than a literal tuple: the list is platform-dependent
# now, and one hardcoded list that silently omitted an engine is what #858
# was. Keep the expectation in a single place, above.
COLI_EXPECT="$SIBLINGS" python3 - <<'PYCHK'
import importlib.machinery, importlib.util, json, os, sys, tempfile
from family_registry import all_families
exe = ".exe" if os.name == "nt" else ""
expect = os.environ["COLI_EXPECT"].split()

Expand All @@ -213,16 +216,25 @@ jobs:
loader.exec_module(cli)

bad = []
for model_type, label, _ in cli._BANNER_MODELS:
with tempfile.TemporaryDirectory() as d:
for family in all_families():
if family.id == "deepseek_v4" and "deepseek_v4" not in expect:
continue
model_type = family.model_types[0]
label = family.display_name
with tempfile.TemporaryDirectory() as d:
with open(os.path.join(d, "config.json"), "w", encoding="utf-8") as f:
json.dump({"model_type": model_type}, f)
open(os.path.join(d, "tokenizer.json"), "w").close()
engine = os.path.basename(cli.engine_for(d))
arch = cli.model_arch(d)
if model_type != "glm" and arch == "glm":
bad.append("%s (%s) -> the GLM engine; the banner names a family "
"the launcher cannot dispatch" % (label, model_type))
expected_file = family.engine_artifact + exe
if not os.path.exists(expected_file):
bad.append("%s: expected artifact %s is absent" %
(label, expected_file))
elif arch != family.id or engine != expected_file:
bad.append("%s (%s) -> %s/%s, expected %s/%s" %
(label, model_type, arch, engine,
family.id, expected_file))
if bad:
sys.exit("FAIL: packaged launcher misroutes:\n " + "\n ".join(bad))
print("OK: every engine is packaged AND coli dispatches to it:", " ".join(expect))
Expand Down
Loading
Loading