Skip to content

[Feature]: Metal backend for kimi_k3 — K3's GPU path is Vulkan-only, so Apple Silicon falls back to CPU #1043

Description

@Toao27

Problem

Kimi K3 is the only tier-1 model with no usable GPU path on macOS, which puts
it out of reach on Apple Silicon regardless of storage.

kimi_k3.c has no Metal support at all:

$ grep -c COLI_METAL kimi_k3.c
0

The Makefile confirms this is by design rather than a build quirk on my end
(Makefile:715):

# kimi_k3.c contains no COLI_CUDA code at all — its GPU path is Vulkan. But

The generic rule still accepts and passes through the Metal flags, so the build
looks like it succeeded when it hasn't:

$ make kimi_k3 METAL=1 ARCH=native
clang ... -mcpu=native -DCOLI_METAL kimi_k3.c   -o kimi_k3 ... -framework Metal -framework Foundation -lc++
kimi_k3.c:255:12: warning: unused variable 'g_k3_vk' [-Wunused-variable]
  255 | static int g_k3_vk=0;     /* backend live (K3_VK=0 disables) */

Two things there: backend_metal.o is absent from the link line — compare the
working GLM build, which includes it —

clang ... -DCOLI_METAL colibri.c  backend_metal.o  -o colibri ... -framework Metal

— and g_k3_vk is unused, so the Vulkan backend didn't compile in either. The
resulting binary is CPU-only with Metal frameworks linked but nothing behind them.

MEASURED IMPACT

I benchmarked Metal vs CPU on GLM-5.2 on an M1 Max (64 GB, 32-core GPU,
commit df2c248) in #1030

RAM_GB=24 Metal 0.49 / CPU 0.20 tok/s = 2.45x
RAM_GB=52 Metal 0.33 / CPU 0.13 tok/s = 2.54x

Prefill is more lopsided still: ~31 s Metal vs ~111 s CPU (3.6x), consistent
with docs/metal.md noting that decode on Apple Silicon is matmul-bound.

Tuned, that machine reaches 0.74 tok/s on GLM-5.2. Without Metal it sits near
0.13 — the difference between usable and not.

For K3 the CPU-only penalty compounds: it routes 16 of 896 experts per layer
against GLM's 8 of 256 (roughly twice the expert bytes per token); at ~1.6 TB
the container won't fit most laptops' internal storage, so it's read over
Thunderbolt (~2.8 GB/s) rather than internal NVMe (4.4 GB/s measured here);
and residency falls accordingly — I see 13% on GLM-5.2 at RAM_GB=24, so K3
would be low single digits.

Extrapolating from the CPU figures, that lands near 0.05 tok/s. A 60-token
answer would take roughly 20 minutes.

Proposed solution

Three options, in rough order of effort — any one of them would help.

  1. A native Metal backend for kimi_k3.c, reusing backend_metal.mm as
    colibri.c does. The kernels GLM needs — batched routed-expert SwiGLU,
    fused decode attention, prefill GEMMs — look like they map onto K3's
    shapes, though K3's larger expert count and different attention may
    need work I can't judge from outside the source.

  2. Failing that, fail loudly. make kimi_k3 METAL=1 currently succeeds and
    produces a binary with no GPU path. A build-time warning, or a startup
    line reading "no GPU backend on this platform", would save people the
    discovery I just made — ideally before they commit to a 1.6 TB download.

  3. Document the per-engine backend matrix, e.g.

    engine CUDA Vulkan Metal
    colibri (GLM-5.2) yes — yes
    kimi_k3 no yes no
    deepseek_v4 ? ? ?
    olmoe ? ? ?

    docs/metal.md describes the Metal backend without saying which engines it
    covers; I assumed it was engine-agnostic until I read the source.

Point 3 alone would have answered my question in seconds and costs almost
nothing.

Alternatives considered

  • Vulkan via MoltenVK. macOS has no native Vulkan driver, so this needs
    MoltenVK translating to Metal underneath plus a working SPIR-V toolchain
    for the VK_SPV shaders. Even if wired up, it adds a translation layer on
    top of the API a Metal path would target directly. A native backend looks
    like the better investment.

  • Accepting CPU-only. At ~0.05 tok/s this isn't a slow-but-usable
    configuration; it's roughly 20 minutes for a short answer, on top of
    €250 of external storage and a 1.6 TB download.

  • Staying on GLM-5.2. This is what I'm doing, and it's a reasonable answer —
    GLM sits within a few index points of K3 on published benchmarks, ships
    under MIT rather than a bespoke licence, and has a working Metal path. But
    it means K3 is effectively unavailable to Apple Silicon users rather than
    merely slow, which seems worth naming.

Scope and compatibility

  • Zero-dependency CPU path: unaffected. This is purely additive, gated behind
    METAL=1, exactly as the GLM Metal path is today. A build without METAL=1
    produces the same binary it does now.
  • CUDA path: unaffected. kimi_k3.c contains no COLI_CUDA code (Makefile:715),
    so there's no interaction.
  • Vulkan path: unaffected — this would be an alternative backend, not a
    replacement. Presumably the same runtime gating that K3_VK uses.
  • Model formats: no change. K3 streams MXFP4 experts natively; this is about
    where the arithmetic runs, not how weights are stored.
  • Platforms: macOS / Apple Silicon only. No effect on Linux, Windows or PowerPC.
  • Options 2 and 3 have no runtime effect at all — a build warning and a docs
    table.

Environment for the observations above: MacBook Pro, M1 Max (10 CPU / 32 GPU
cores), 64 GiB unified memory, macOS 26.6.1 (25G76), Apple clang 21.0.0,
Homebrew libomp 22.1.8, commit df2c248
(v1.6.1, tip of main). The GLM-5.2 Metal build works correctly on this machine.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    featureNuova funzionalitàmetalBackend Metal/Applemodel-supportSupporto a nuovi modelli

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions