Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -68,3 +68,12 @@ olmoe_merged/
olmoe_i4/
c/olmoe_merged/
c/olmoe_i4/

# colibri local weights + vulkan SDK (large, not source)
/modles/
/c/vulkan/
/c/tools/glslang/
/c/tools/glslang.zip
/c/tools/vkheaders/
/c/tools/vk_headers.zip
/c/tools/vk_tmp/
33 changes: 31 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -253,6 +253,34 @@ the full 756 GB on disk at once:
./coli convert --model /nvme/glm52_i4 # download+convert shard by shard (python, one-time)
```

#### Qwen3.6-35B-A3B (35B / 3B active) — community build

A pure-C **Qwen3.6-35B-A3B** runtime (`c/qwen36.c`) is available as a community
contribution: Gated-Attention (GQA + partial RoPE) + Gated DeltaNet recurrent
linear attention, with a streaming MoE, an optional **Vulkan** MoE backend
(int4 unpacked in-shader + float GEMV, works around the AMD integrated-GPU
`OpSDotKHR` segfault) and resident-expert pinning (`COLIBRI_RESIDENT=1/2`) to
cut disk IO. Runs on a 16 GB laptop.

Pre-converted colibri containers are on Hugging Face:

- **int8** (~35 GB, lower quant loss): <https://huggingface.co/minne100/qwen36-35b-a3b-colibri-i8>
- **int4** (~20 GB): <https://huggingface.co/minne100/qwen36-35b-a3b-colibri-i4>
- conversion toolkit + reference prompts: <https://huggingface.co/datasets/minne100/colibri-qwen36-tools>

```bash
# build
cd c && gcc -D_FILE_OFFSET_BITS=64 -O3 -march=native -fopenmp -o qwen36.exe qwen36.c vulkan_gemv.c -lm -fopenmp -static -lpsapi
# run (uses Vulkan GPU if present, otherwise falls back to CPU)
SNAP=<model> TOK=<model>/tokenizer.json ./qwen36 8 4 prompt.txt
```

> Built in PR [#602](https://github.com/JustVugg/colibri/pull/602). On an
> integrated-GPU 16 GB machine a cold-start benchmark shows **int8 CPU-only is
> fastest** (~1.08 tok/s, 10.25 GB peak) because the iGPU shares system memory
> and gains no bandwidth, while int4 pays an unpack penalty; int4 stays half
> the on-disk size.

### 3. Run it

```bash
Expand Down Expand Up @@ -289,8 +317,9 @@ and the optional API gateway.
merged in the open.
- **More open models.** The tiering algorithm is model-agnostic: any MoE with
routed experts can be staged the same way. GLM-5.2 and OLMoE run today;
support for more open-weight families — **Kimi K2** (Moonshot AI),
**Qwen3 MoE** (Alibaba), **MiniMax** — is on the roadmap.
**Qwen3.6-35B-A3B** is now available as a community build (PR
[#602](https://github.com/JustVugg/colibri/pull/602)); more open-weight
families — **Kimi K2** (Moonshot AI), **MiniMax** — are on the roadmap.

## Supporting the project

Expand Down
12 changes: 11 additions & 1 deletion c/Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -333,6 +333,15 @@ cuda-bench: backend_cuda.cu backend_cuda.h backend_gpu_compat.h tests/bench_tens
olmoe$(EXE): olmoe.c st.h json.h compat.h
$(CC) $(CFLAGS) olmoe.c -o olmoe$(EXE) $(LDFLAGS)

# Qwen3.6-35B-A3B Phase-2 engine: Gated Attention + Gated DeltaNet (recurrent
# linear attention) + streaming MoE. Validates against tools/_ref_dn.py (numpy
# vs HF oracle) and via DUMP logit-cosine against tools/_ref_dn.py --dump.
qwen36$(EXE): qwen36.c vulkan_gemv.c st.h json.h compat.h vulkan_gemv.h vulkan_core.h vulkan_gemv_spv.h
$(CC) $(CFLAGS) qwen36.c vulkan_gemv.c -o qwen36$(EXE) $(LDFLAGS)

qwen36_serve$(EXE): qwen36_serve.c vulkan_gemv.c st.h json.h compat.h vulkan_gemv.h vulkan_core.h vulkan_gemv_spv.h
$(CC) $(CFLAGS) qwen36_serve.c vulkan_gemv.c -o qwen36_serve$(EXE) $(LDFLAGS) -lws2_32 -Wl,--stack,16777216

# Use a baseline that matches the compiler target. macOS already targets a
# portable baseline when ARCH is empty; forcing the x86 value there breaks
# Apple Silicon. Unknown targets use native rather than an invalid x86 flag.
Expand Down Expand Up @@ -488,13 +497,14 @@ check:
$(MAKE) portable
$(MAKE) test

install: colibri$(EXE) olmoe$(EXE)
install: colibri$(EXE) olmoe$(EXE) qwen36$(EXE)
$(INSTALL) -d $(DESTDIR)$(BINDIR)
$(INSTALL) -d $(DESTDIR)$(LIBEXECDIR)
$(INSTALL) -d $(DESTDIR)$(LIBEXECDIR)/tools
$(INSTALL) -m 755 coli $(DESTDIR)$(BINDIR)/coli
$(INSTALL) -m 755 colibri$(EXE) $(DESTDIR)$(LIBEXECDIR)/colibri$(EXE)
$(INSTALL) -m 755 olmoe$(EXE) $(DESTDIR)$(LIBEXECDIR)/olmoe$(EXE)
$(INSTALL) -m 755 qwen36$(EXE) $(DESTDIR)$(LIBEXECDIR)/qwen36$(EXE)
$(INSTALL) -m 644 resource_plan.py doctor.py openai_server.py $(DESTDIR)$(LIBEXECDIR)/
$(INSTALL) -m 644 tools/*.py $(DESTDIR)$(LIBEXECDIR)/tools/

Expand Down
42 changes: 42 additions & 0 deletions c/int4_vs_int8_cold.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
# int4 / int8(+GPU) / int8(CPU) 冷启动对比 — Qwen3.6-35B-A3B

**测试时间**:2026-07-25
**引擎**:`c/qwen36.exe`(纯 C,长问题 `c/prompt_emerge.txt` *"详细介绍一下LLM智能涌现的原理,不少于1000字。"*)
**运行方式**:命令行冷启动(非 serve),串行各跑一次,`N_NEW=256`
**硬件**:本机 16GB RAM,AMD 780M 核显(Vulkan,**共享系统内存、无独立显存**)
**命令**:`./qwen36.exe 8 <bits> C:/.../prompt_emerge.txt`,`SNAP`/`TOK` 指向对应精度目录;int8-CPU 额外 `COLIBRI_GPU=0`
**GPU 路径**:int4 = in-shader unpack+float GEMV(480MB 上传);int8 = float path(960MB 上传,因 AMD 驱动 0x800184 的 OpSDotKHR bug 禁用 int8 dot-product)

## 对比表

| 指标 | int4 +GPU | int8 +GPU | int8 CPU-only |
|---|---|---|---|
| **首字延迟 TTFT** | 50.18 s | 28.68 s | **13.83 s** |
| **每秒速度** | 0.29 tok/s | 0.48 tok/s | **1.08 tok/s** |
| **PEAK RSS(内存)** | 11.30 GB | 11.27 GB | **10.25 GB** |
| 加载后 RSS | 9.25 GB | 9.25 GB | 9.24 GB |
| 专家缓存命中率 | 30.4% | 30.4% | 30.4% |
| 256 token 总耗时 | 873.2 s | 532.0 s | **238.0 s** |
| 生成内容 | 正确 | 正确 | 正确(细节措辞略异)|

## 关键发现

1. **内存**:加载后三者均 ~9.25GB;峰值 int8-CPU 最小(10.25GB,无 GPU 上传副本),int4/int8+GPU 因核显共享内存上传权重副本各 ~11.3GB。int8 未 OOM(16GB 充裕 + `cap=8` LRU 未填满)。

2. **速度排序:int8-CPU > int8+GPU > int4+GPU**。
- int4+GPU 最慢(0.29):int4 每专家需在 CPU 侧 unpack→int8 再传 GPU(日志 `"unpacking to int8 in slot"`),核显又无带宽优势,双重拖累。
- int8+GPU 快于 int4+GPU(0.48):走 GPU float path 免解包,但仍受共享内存上传开销。
- int8-CPU 最快(1.08):直接 CPU 算 int8,无 unpack、无 GPU 上传,反而最干净。

3. **TTFT**:int8-CPU 13.83s 最低(prefill 直接 CPU 算,无上传);int4+GPU 50.18s 最高(prefill 也受 unpack 拖累)。

4. **生成内容**:三者主题一致、质量相当,均正确阐述 LLM 涌现的规模效应/相变理论。int8-CPU 细节措辞与 GPU 版略有差异(如"发生了非线性的跃迁" vs "呈现出非线性的跳跃"),到 256 token 截断点不同,但无质量退化。

## 结论

本机**核显(共享内存、无独显)冷启动**下真实排序:**int8 CPU-only 最优**(速度 3.7× int4-GPU、内存最小),其次 int8+GPU,最后 int4+GPU。GPU 在核显上不仅没加速,反而因 int4 解包 + 共享内存上传变慢。

**int4 唯一硬优势是磁盘占用小**(19.7GB vs 34.6GB)——仅当磁盘吃紧才选 int4。
**独显 + 常驻显存**机器上 GPU 优势才会真正显现(独立高带宽显存),该场景未测。

> 注:为测首字延迟,给 `qwen36.c` 非 OpenAI 文本模式加了 TTFT stderr 打印;bash 下 prompt 参数须用 `C:/...` 大写盘符(MinGW fopen 不认 `/c/`)。
1 change: 1 addition & 0 deletions c/prompt_emerge.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
详细介绍一下LLM智能涌现的原理,不少于1000字。
Loading
Loading