Problem
Kimi K3 is the only tier-1 model with no usable GPU path on macOS, which puts
it out of reach on Apple Silicon regardless of storage.
kimi_k3.c has no Metal support at all:
$ grep -c COLI_METAL kimi_k3.c
0
The Makefile confirms this is by design rather than a build quirk on my end
(Makefile:715):
# kimi_k3.c contains no COLI_CUDA code at all — its GPU path is Vulkan. But
The generic rule still accepts and passes through the Metal flags, so the build
looks like it succeeded when it hasn't:
$ make kimi_k3 METAL=1 ARCH=native
clang ... -mcpu=native -DCOLI_METAL kimi_k3.c -o kimi_k3 ... -framework Metal -framework Foundation -lc++
kimi_k3.c:255:12: warning: unused variable 'g_k3_vk' [-Wunused-variable]
255 | static int g_k3_vk=0; /* backend live (K3_VK=0 disables) */
Two things there: backend_metal.o is absent from the link line — compare the
working GLM build, which includes it —
clang ... -DCOLI_METAL colibri.c backend_metal.o -o colibri ... -framework Metal
— and g_k3_vk is unused, so the Vulkan backend didn't compile in either. The
resulting binary is CPU-only with Metal frameworks linked but nothing behind them.
MEASURED IMPACT
I benchmarked Metal vs CPU on GLM-5.2 on an M1 Max (64 GB, 32-core GPU,
commit df2c248) in #1030
RAM_GB=24 Metal 0.49 / CPU 0.20 tok/s = 2.45x
RAM_GB=52 Metal 0.33 / CPU 0.13 tok/s = 2.54x
Prefill is more lopsided still: ~31 s Metal vs ~111 s CPU (3.6x), consistent
with docs/metal.md noting that decode on Apple Silicon is matmul-bound.
Tuned, that machine reaches 0.74 tok/s on GLM-5.2. Without Metal it sits near
0.13 — the difference between usable and not.
For K3 the CPU-only penalty compounds: it routes 16 of 896 experts per layer
against GLM's 8 of 256 (roughly twice the expert bytes per token); at ~1.6 TB
the container won't fit most laptops' internal storage, so it's read over
Thunderbolt (~2.8 GB/s) rather than internal NVMe (4.4 GB/s measured here);
and residency falls accordingly — I see 13% on GLM-5.2 at RAM_GB=24, so K3
would be low single digits.
Extrapolating from the CPU figures, that lands near 0.05 tok/s. A 60-token
answer would take roughly 20 minutes.
Proposed solution
Three options, in rough order of effort — any one of them would help.
-
A native Metal backend for kimi_k3.c, reusing backend_metal.mm as
colibri.c does. The kernels GLM needs — batched routed-expert SwiGLU,
fused decode attention, prefill GEMMs — look like they map onto K3's
shapes, though K3's larger expert count and different attention may
need work I can't judge from outside the source.
-
Failing that, fail loudly. make kimi_k3 METAL=1 currently succeeds and
produces a binary with no GPU path. A build-time warning, or a startup
line reading "no GPU backend on this platform", would save people the
discovery I just made — ideally before they commit to a 1.6 TB download.
-
Document the per-engine backend matrix, e.g.
| engine |
CUDA |
Vulkan |
Metal |
| colibri (GLM-5.2) |
yes |
— |
yes |
| kimi_k3 |
no |
yes |
no |
| deepseek_v4 |
? |
? |
? |
| olmoe |
? |
? |
? |
docs/metal.md describes the Metal backend without saying which engines it
covers; I assumed it was engine-agnostic until I read the source.
Point 3 alone would have answered my question in seconds and costs almost
nothing.
Alternatives considered
-
Vulkan via MoltenVK. macOS has no native Vulkan driver, so this needs
MoltenVK translating to Metal underneath plus a working SPIR-V toolchain
for the VK_SPV shaders. Even if wired up, it adds a translation layer on
top of the API a Metal path would target directly. A native backend looks
like the better investment.
-
Accepting CPU-only. At ~0.05 tok/s this isn't a slow-but-usable
configuration; it's roughly 20 minutes for a short answer, on top of
€250 of external storage and a 1.6 TB download.
-
Staying on GLM-5.2. This is what I'm doing, and it's a reasonable answer —
GLM sits within a few index points of K3 on published benchmarks, ships
under MIT rather than a bespoke licence, and has a working Metal path. But
it means K3 is effectively unavailable to Apple Silicon users rather than
merely slow, which seems worth naming.
Scope and compatibility
- Zero-dependency CPU path: unaffected. This is purely additive, gated behind
METAL=1, exactly as the GLM Metal path is today. A build without METAL=1
produces the same binary it does now.
- CUDA path: unaffected. kimi_k3.c contains no COLI_CUDA code (Makefile:715),
so there's no interaction.
- Vulkan path: unaffected — this would be an alternative backend, not a
replacement. Presumably the same runtime gating that K3_VK uses.
- Model formats: no change. K3 streams MXFP4 experts natively; this is about
where the arithmetic runs, not how weights are stored.
- Platforms: macOS / Apple Silicon only. No effect on Linux, Windows or PowerPC.
- Options 2 and 3 have no runtime effect at all — a build warning and a docs
table.
Environment for the observations above: MacBook Pro, M1 Max (10 CPU / 32 GPU
cores), 64 GiB unified memory, macOS 26.6.1 (25G76), Apple clang 21.0.0,
Homebrew libomp 22.1.8, commit df2c248
(v1.6.1, tip of main). The GLM-5.2 Metal build works correctly on this machine.
Problem
Kimi K3 is the only tier-1 model with no usable GPU path on macOS, which puts
it out of reach on Apple Silicon regardless of storage.
kimi_k3.chas no Metal support at all:The Makefile confirms this is by design rather than a build quirk on my end
(Makefile:715):
The generic rule still accepts and passes through the Metal flags, so the build
looks like it succeeded when it hasn't:
Two things there:
backend_metal.ois absent from the link line — compare theworking GLM build, which includes it —
— and
g_k3_vkis unused, so the Vulkan backend didn't compile in either. Theresulting binary is CPU-only with Metal frameworks linked but nothing behind them.
MEASURED IMPACT
I benchmarked Metal vs CPU on GLM-5.2 on an M1 Max (64 GB, 32-core GPU,
commit df2c248) in #1030
RAM_GB=24 Metal 0.49 / CPU 0.20 tok/s = 2.45x
RAM_GB=52 Metal 0.33 / CPU 0.13 tok/s = 2.54x
Prefill is more lopsided still: ~31 s Metal vs ~111 s CPU (3.6x), consistent
with docs/metal.md noting that decode on Apple Silicon is matmul-bound.
Tuned, that machine reaches 0.74 tok/s on GLM-5.2. Without Metal it sits near
0.13 — the difference between usable and not.
For K3 the CPU-only penalty compounds: it routes 16 of 896 experts per layer
against GLM's 8 of 256 (roughly twice the expert bytes per token); at ~1.6 TB
the container won't fit most laptops' internal storage, so it's read over
Thunderbolt (~2.8 GB/s) rather than internal NVMe (4.4 GB/s measured here);
and residency falls accordingly — I see 13% on GLM-5.2 at RAM_GB=24, so K3
would be low single digits.
Extrapolating from the CPU figures, that lands near 0.05 tok/s. A 60-token
answer would take roughly 20 minutes.
Proposed solution
Three options, in rough order of effort — any one of them would help.
A native Metal backend for kimi_k3.c, reusing backend_metal.mm as
colibri.c does. The kernels GLM needs — batched routed-expert SwiGLU,
fused decode attention, prefill GEMMs — look like they map onto K3's
shapes, though K3's larger expert count and different attention may
need work I can't judge from outside the source.
Failing that, fail loudly.
make kimi_k3 METAL=1currently succeeds andproduces a binary with no GPU path. A build-time warning, or a startup
line reading "no GPU backend on this platform", would save people the
discovery I just made — ideally before they commit to a 1.6 TB download.
Document the per-engine backend matrix, e.g.
docs/metal.md describes the Metal backend without saying which engines it
covers; I assumed it was engine-agnostic until I read the source.
Point 3 alone would have answered my question in seconds and costs almost
nothing.
Alternatives considered
Vulkan via MoltenVK. macOS has no native Vulkan driver, so this needs
MoltenVK translating to Metal underneath plus a working SPIR-V toolchain
for the VK_SPV shaders. Even if wired up, it adds a translation layer on
top of the API a Metal path would target directly. A native backend looks
like the better investment.
Accepting CPU-only. At ~0.05 tok/s this isn't a slow-but-usable
configuration; it's roughly 20 minutes for a short answer, on top of
€250 of external storage and a 1.6 TB download.
Staying on GLM-5.2. This is what I'm doing, and it's a reasonable answer —
GLM sits within a few index points of K3 on published benchmarks, ships
under MIT rather than a bespoke licence, and has a working Metal path. But
it means K3 is effectively unavailable to Apple Silicon users rather than
merely slow, which seems worth naming.
Scope and compatibility
METAL=1, exactly as the GLM Metal path is today. A build without METAL=1
produces the same binary it does now.
so there's no interaction.
replacement. Presumably the same runtime gating that K3_VK uses.
where the arithmetic runs, not how weights are stored.
table.
Environment for the observations above: MacBook Pro, M1 Max (10 CPU / 32 GPU
cores), 64 GiB unified memory, macOS 26.6.1 (25G76), Apple clang 21.0.0,
Homebrew libomp 22.1.8, commit df2c248
(v1.6.1, tip of main). The GLM-5.2 Metal build works correctly on this machine.