Orchestration + dashboard for self-hosted AI inference. The control plane that lets you stop SSH-ing into your GPU box to swap models — point your apps at one URL and let Witchgrid decide what runs where, right now.
For the case where you have a small fleet of CPU/GPU boxes and want to manage what's running where, what's loaded, and what's available — without hand-wiring each consumer to a hardcoded inference URL.
Status: v0.9.3 — 1.0 candidate. Single-operator, LAN-first. In daily production use serving flammen.ai. Multi-node capacity-aware placement, auto-spawn, on-demand model management, a real-time dashboard (SSE-pushed, failure-first), a self-healing service supervisor, a first-class CPU/GPU toggle, a guided profile wizard, and graceful WAL-checkpoint shutdown all work. The whole 1.0 punch list is checked; the 0.9.x line is the soak-before-1.0 cut.
Built in Hemlock (2.8.2). Each binary is a single ~10 MB ELF. Since Hemlock 2.6.0 hemlockc static-links libwebsockets/libssl/libffi, so they are build-time inputs, not runtime deps; what's left dynamic is universal (libcap, libuv, libev, libm, libc). The one real runtime requirement is libsqlite3 — @stdlib/sqlite dlopen()s libsqlite3.so.0 at startup, so it never shows up in ldd and must be present on every box that runs the CP or an agent. Deploy is scp + a service unit.
- Node registry + heartbeats — every box runs an agent that reports hardware (GPU model, total/free VRAM via
nvidia-smi, CPU, RAM, disk) on register and every 30 s. Stale (> 90 s) nodes drop out of placement. - Capacity-aware placement — a GGUF metadata parser extracts model weights + KV-cache footprint at decision time; placement picks the (node, GPU set) that fits in live free VRAM. Tensor-split across GPUs when needed.
- Auto-spawn on demand —
POST /v1/llama/{profile}/*proxies to a running instance, or — if none — places + spawns one, waits for its/health, then completes the request. Spawn nothing by hand. - Self-healing supervision — the agent keeps its services as desired state: an auto-restart watchdog (default on, per-node toggle) respawns any service that dies, covering both crashes and host reboots, with crash-loop backoff. Orphan adoption re-attaches services that outlived an agent crash.
- Multi-engine — not just llama.cpp. Manages
llama-server(chat/embeddings/vision),sd-server(stable-diffusion.cpp images),whisper-server(STT), andpiper(TTS), each with a normalized profile + on-demand install/activate. - Model management — agents scan model dirs into a per-node catalog (advertised by alias, so profiles stop pinning host paths); CP can pull from HuggingFace and distribute models peer-to-peer between nodes.
- Intent-driven argv — profiles declare normalized intents (
flash_attention: true,kv_cache_k: "q4_0",context: 131072); the agent translates each into the binary's actual flag form from its parsed--help, so one profile produces the right argv across llama.cpp versions. - Routing + SSE — direct
/resolve/{profile}for resolve-then-connect consumers (CP stays off the data path), or the/v1/llamaproxy with SSE passthrough (stream:truestreamstext/event-streamstraight through). - Observability — Prometheus
/metrics(node liveness, per-GPU VRAM, watchdog state) + a live HTMX dashboard (nodes, services, profiles, capabilities, catalog, transfers, bench). - MCP server — the CP speaks the Model Context Protocol over Streamable HTTP at
POST /mcp, exposing the fleet as tools (list_nodes,list_services,list_profiles,list_catalog,fleet_status,resolve_profile,spawn_service,stop_service). Point Claude (or any MCP client) at it to introspect and drive Witchgrid in natural language. Gated by the same shared bearer as the rest of the operator API.
- Not a job queue. Consumers bring their own (pgmq, RabbitMQ, …). Witchgrid answers "where does this request go now," not "when does this job run."
- Not a fine-tuning harness. See Merlina — the intended composition is Merlina trains, Witchgrid serves.
- Not Kubernetes. Single-binary CP + lightweight agents over plain HTTP. No pods, no manifests, no networking abstractions.
- Not multi-tenant (yet). Single operator. Auth is an optional shared bearer secret (
WITCHGRID_SHARED_SECRET), with opt-in knobs to also gate the read-only surface (WITCHGRID_AUTH_PROTECT_READ) and the/v1/llama/*inference plane (WITCHGRID_AUTH_PROTECT_INFERENCE) — seedocs/security.md. The design target is a trusted LAN — don't expose the CP to the internet. - Not for single-node setups. One machine? Just SSH and edit your scripts. The value kicks in at ≥2 nodes, or when sharing a GPU between workloads.
┌──────────────┐ embedded HTMX dashboard + Prometheus /metrics
│ operator / │
│ Grafana │
└──────┬───────┘
▼
┌──────────────────┐ ┌───────────────────────────────┐
│ control plane │ ◄───── │ consumers │
│ cp/ :8765 │ HTTP │ resolve-then-connect, or the │
│ registry · place │ │ /v1/llama proxy (+ SSE) │
│ · route · catalog│ └───────────────────────────────┘
└────────┬─────────┘
│ /register heartbeats + spawn/stop/settings RPC
┌─────┼───────────────┬────────────────────┐
▼ ▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────────┐
│ agent │ │ agent │ │ agent (cpu) │
│ :8766 │ │ :8766 │ │ :8766 │
│ llama / │ │ sd / │ │ whisper / │
│ sd /... │ │ llama │ │ piper │
│ :18xxx │ │ :19xxx │ │ :19xxx │
└──────────┘ └──────────┘ └──────────────┘
cp/ and agent/ each compile to a standalone ELF (witchgrid-cp, witchgrid-agent). Profiles, the node registry, and routing all live on the CP; agents are stateless-ish supervisors that run whatever the CP inlines into a spawn.
# Control plane — on the box you'll call from
cd cp && hemlockc cp.hml -o witchgrid-cp && ./witchgrid-cp
# → :8765, dashboard at GET /, metrics at GET /metrics
# Agent — on every box that hosts inference
cd agent && hemlockc agent.hml -o witchgrid-agent
WITCHGRID_CP_URL=http://<cp-host>:8765 \
WITCHGRID_AGENT_URL=http://<this-host>:8766 \
WITCHGRID_DATA_DIR=$HOME/.witchgrid \
./witchgrid-agent
# → registers, heartbeats every 30s, supervises engines on demand
# Drive it
curl -s http://<cp-host>:8765/nodes | jq # who's in the grid
curl -X POST http://<cp-host>:8765/v1/llama/chat-llama3.1-8b/completion \
-H 'content-type: application/json' \
-d '{"prompt":"hello","n_predict":32}' # auto-place + spawn + completeA fresh CP seeds a set of ready-to-use default profiles — no config to write before the first request. General-purpose starters (all auto-pull their GGUF from a trusted quanter on first spawn):
- Chat by VRAM tier —
chat-llama3.2-3b(CPU/8 GB),chat-llama3.1-8b(12-16 GB),chat-qwen2.5-14b(24 GB),chat-qwen2.5-32b(32-48 GB). - Trade-offs on the 8B workhorse —
chat-llama3.1-8b-quality(Q6_K, fidelity over speed) andchat-llama3.1-8b-serving(4 slots, throughput over latency). - Specialists —
reason-deepseek-r1-14b(reasoning),code-qwen2.5-7b/code-qwen2.5-32b(coding),embed-nomic(embeddings),vision-qwen(multimodal),audio-whisper-base(STT),tts-piper-en(TTS).
Plus the flammen.ai production profiles (chat-mahou, designer-cpu, memory-phoenix, image-niku) kept as worked examples. Clone any of them in the dashboard's profile wizard to retune for your hardware. Defaults are filled in with insert-or-ignore, so upgrades pick up new ones without touching profiles you've edited.
Prebuilt binaries (linux-x86_64 + macos-arm64) are attached to each GitHub release; installing the matching Hemlock runtime is only needed to build.
| Doc | What's in it |
|---|---|
docs/operations.md |
Deploy + run in production: service units (systemd / pm2 / launchd), the full env-var reference, auto-restart, metrics scraping, backup/restore, reboot recovery, upgrades. |
cp/README.md |
Control-plane endpoint reference + build/run + smoke test. |
agent/README.md |
Agent endpoints, config, profiles (owned by CP), engine install lifecycle, restart behavior. |
CLAUDE.md |
Architecture, components, design decisions, project relationships. |
docs/test-cases-todo.md |
Test backlog (what's automated vs. still manual). |
docs/llama-cpp-as-managed-runtime.md |
The "Witchgrid as the llama.cpp manager" long-game (vendoring, canary/rollback). |
| Dir | What's there |
|---|---|
cp/ |
Control-plane daemon: registry, placement, routing, profiles, catalog, dashboard, GGUF introspection. |
agent/ |
Per-node agent: hardware probe, capability introspection, engine supervisor + watchdog, install lifecycle. |
dashboard/ |
HTMX + Alpine + Pico assets, embedded into cp via scripts/embed_assets.sh. |
scripts/ |
Build helpers (embed_assets.sh) + the CP systemd unit template. |
tests/ |
Unit suite (hemlock tests/run.hml) + integration harness (tests/integration/). |
docs/ |
Design docs, ops guide, roadmaps, the Hemlock language-eval archive. |
legacy-python/ |
Frozen v0.3 Python prototype. Reference only. |
hemlock tests/run.hml # unit: intent / capabilities / GGUF parser
bash tests/integration/run_integration.sh # integration: real CP+agent, fake backendCI (.github/workflows/build.yml) builds both platforms, runs the unit + integration suites, and on a v* tag attaches per-platform binaries + SHA256SUMS to the release. See tests/README.md.
Capability-wise it's there; 1.0 is about coherence, deployability, and test breadth:
- Single source of truth for the version (
version.hml), surfaced in/healthz+/metrics(witchgrid_build_info) + the dashboard header. - CP bind configurable —
WITCHGRID_CP_HOST/WITCHGRID_CP_PORT. - Deploy/ops documented — see
docs/operations.md. - Test breadth — the integration harness now covers port-contention, stale-node drop, spawn-failure, bad-model-path, multi-node registration, the CPU/GPU device override, and the live/health endpoints (26 assertions).
- Auth posture finalized — LAN-first shared-secret model, documented in
docs/security.md(what's public vs protected, reverse-proxy guidance for exposure). Deferred to post-1.0: a dashboard login (slice 2) + per-consumer/v1/llamatokens (slice 3) + native TLS. - Graceful shutdown — on SIGINT/TERM/HUP the CP checkpoints the WAL (TRUNCATE) and exits cleanly. Live as of the Hemlock 2.6.2 pin, which carries the async-runtime signal fix (hemlang/hemlock#587). Single-host multi-node registration is tested; a cross-node 2-agent test (mocked
nvidia-smi) is the remaining test nicety.
- Merlina — the lab's fine-tuning UI. Merlina trains, Witchgrid serves.
- flammen.ai — Witchgrid's first consumer (chat + image + caption workers route through it).
Built in Hemlock. The language-evaluation writeup that informed the rewrite lives in docs/hemlock-feedback-archive.md.