Skip to content

Latest commit

 

History

132 Commits

Folders and files

Repository files navigation

Witchgrid

Orchestration + dashboard for self-hosted AI inference. The control plane that lets you stop SSH-ing into your GPU box to swap models — point your apps at one URL and let Witchgrid decide what runs where, right now.

For the case where you have a small fleet of CPU/GPU boxes and want to manage what's running where, what's loaded, and what's available — without hand-wiring each consumer to a hardcoded inference URL.

Status: v0.9.3 — 1.0 candidate. Single-operator, LAN-first. In daily production use serving flammen.ai. Multi-node capacity-aware placement, auto-spawn, on-demand model management, a real-time dashboard (SSE-pushed, failure-first), a self-healing service supervisor, a first-class CPU/GPU toggle, a guided profile wizard, and graceful WAL-checkpoint shutdown all work. The whole 1.0 punch list is checked; the 0.9.x line is the soak-before-1.0 cut.

Built in Hemlock (2.8.2). Each binary is a single ~10 MB ELF. Since Hemlock 2.6.0 hemlockc static-links libwebsockets/libssl/libffi, so they are build-time inputs, not runtime deps; what's left dynamic is universal (libcap, libuv, libev, libm, libc). The one real runtime requirement is libsqlite3 — @stdlib/sqlite dlopen()s libsqlite3.so.0 at startup, so it never shows up in ldd and must be present on every box that runs the CP or an agent. Deploy is scp + a service unit.


What it does

  • Node registry + heartbeats — every box runs an agent that reports hardware (GPU model, total/free VRAM via nvidia-smi, CPU, RAM, disk) on register and every 30 s. Stale (> 90 s) nodes drop out of placement.
  • Capacity-aware placement — a GGUF metadata parser extracts model weights + KV-cache footprint at decision time; placement picks the (node, GPU set) that fits in live free VRAM. Tensor-split across GPUs when needed.
  • Auto-spawn on demand — POST /v1/llama/{profile}/* proxies to a running instance, or — if none — places + spawns one, waits for its /health, then completes the request. Spawn nothing by hand.
  • Self-healing supervision — the agent keeps its services as desired state: an auto-restart watchdog (default on, per-node toggle) respawns any service that dies, covering both crashes and host reboots, with crash-loop backoff. Orphan adoption re-attaches services that outlived an agent crash.
  • Multi-engine — not just llama.cpp. Manages llama-server (chat/embeddings/vision), sd-server (stable-diffusion.cpp images), whisper-server (STT), and piper (TTS), each with a normalized profile + on-demand install/activate.
  • Model management — agents scan model dirs into a per-node catalog (advertised by alias, so profiles stop pinning host paths); CP can pull from HuggingFace and distribute models peer-to-peer between nodes.
  • Intent-driven argv — profiles declare normalized intents (flash_attention: true, kv_cache_k: "q4_0", context: 131072); the agent translates each into the binary's actual flag form from its parsed --help, so one profile produces the right argv across llama.cpp versions.
  • Routing + SSE — direct /resolve/{profile} for resolve-then-connect consumers (CP stays off the data path), or the /v1/llama proxy with SSE passthrough (stream:true streams text/event-stream straight through).
  • Observability — Prometheus /metrics (node liveness, per-GPU VRAM, watchdog state) + a live HTMX dashboard (nodes, services, profiles, capabilities, catalog, transfers, bench).
  • MCP server — the CP speaks the Model Context Protocol over Streamable HTTP at POST /mcp, exposing the fleet as tools (list_nodes, list_services, list_profiles, list_catalog, fleet_status, resolve_profile, spawn_service, stop_service). Point Claude (or any MCP client) at it to introspect and drive Witchgrid in natural language. Gated by the same shared bearer as the rest of the operator API.

What it isn't

  • Not a job queue. Consumers bring their own (pgmq, RabbitMQ, …). Witchgrid answers "where does this request go now," not "when does this job run."
  • Not a fine-tuning harness. See Merlina — the intended composition is Merlina trains, Witchgrid serves.
  • Not Kubernetes. Single-binary CP + lightweight agents over plain HTTP. No pods, no manifests, no networking abstractions.
  • Not multi-tenant (yet). Single operator. Auth is an optional shared bearer secret (WITCHGRID_SHARED_SECRET), with opt-in knobs to also gate the read-only surface (WITCHGRID_AUTH_PROTECT_READ) and the /v1/llama/* inference plane (WITCHGRID_AUTH_PROTECT_INFERENCE) — see docs/security.md. The design target is a trusted LAN — don't expose the CP to the internet.
  • Not for single-node setups. One machine? Just SSH and edit your scripts. The value kicks in at ≥2 nodes, or when sharing a GPU between workloads.

Architecture

┌──────────────┐   embedded HTMX dashboard + Prometheus /metrics
│  operator /  │
│  Grafana     │
└──────┬───────┘
       ▼
┌──────────────────┐        ┌───────────────────────────────┐
│  control plane   │ ◄───── │  consumers                    │
│   cp/  :8765     │  HTTP  │  resolve-then-connect, or the │
│  registry · place │       │  /v1/llama proxy (+ SSE)      │
│   · route · catalog│      └───────────────────────────────┘
└────────┬─────────┘
         │  /register heartbeats  +  spawn/stop/settings RPC
   ┌─────┼───────────────┬────────────────────┐
   ▼     ▼               ▼                    ▼
┌──────────┐      ┌──────────┐        ┌──────────────┐
│ agent    │      │ agent    │        │ agent (cpu)  │
│  :8766   │      │  :8766   │        │  :8766       │
│ llama /  │      │ sd /     │        │ whisper /    │
│ sd /...  │      │ llama    │        │ piper        │
│  :18xxx  │      │  :19xxx  │        │  :19xxx      │
└──────────┘      └──────────┘        └──────────────┘

cp/ and agent/ each compile to a standalone ELF (witchgrid-cp, witchgrid-agent). Profiles, the node registry, and routing all live on the CP; agents are stateless-ish supervisors that run whatever the CP inlines into a spawn.

Quickstart

# Control plane — on the box you'll call from
cd cp && hemlockc cp.hml -o witchgrid-cp && ./witchgrid-cp
#  → :8765, dashboard at GET /, metrics at GET /metrics

# Agent — on every box that hosts inference
cd agent && hemlockc agent.hml -o witchgrid-agent
WITCHGRID_CP_URL=http://<cp-host>:8765 \
WITCHGRID_AGENT_URL=http://<this-host>:8766 \
WITCHGRID_DATA_DIR=$HOME/.witchgrid \
  ./witchgrid-agent
#  → registers, heartbeats every 30s, supervises engines on demand

# Drive it
curl -s http://<cp-host>:8765/nodes | jq            # who's in the grid
curl -X POST http://<cp-host>:8765/v1/llama/chat-llama3.1-8b/completion \
  -H 'content-type: application/json' \
  -d '{"prompt":"hello","n_predict":32}'             # auto-place + spawn + complete

A fresh CP seeds a set of ready-to-use default profiles — no config to write before the first request. General-purpose starters (all auto-pull their GGUF from a trusted quanter on first spawn):

  • Chat by VRAM tier — chat-llama3.2-3b (CPU/8 GB), chat-llama3.1-8b (12-16 GB), chat-qwen2.5-14b (24 GB), chat-qwen2.5-32b (32-48 GB).
  • Trade-offs on the 8B workhorse — chat-llama3.1-8b-quality (Q6_K, fidelity over speed) and chat-llama3.1-8b-serving (4 slots, throughput over latency).
  • Specialists — reason-deepseek-r1-14b (reasoning), code-qwen2.5-7b / code-qwen2.5-32b (coding), embed-nomic (embeddings), vision-qwen (multimodal), audio-whisper-base (STT), tts-piper-en (TTS).

Plus the flammen.ai production profiles (chat-mahou, designer-cpu, memory-phoenix, image-niku) kept as worked examples. Clone any of them in the dashboard's profile wizard to retune for your hardware. Defaults are filled in with insert-or-ignore, so upgrades pick up new ones without touching profiles you've edited.

Prebuilt binaries (linux-x86_64 + macos-arm64) are attached to each GitHub release; installing the matching Hemlock runtime is only needed to build.

Docs

Doc What's in it
docs/operations.md Deploy + run in production: service units (systemd / pm2 / launchd), the full env-var reference, auto-restart, metrics scraping, backup/restore, reboot recovery, upgrades.
cp/README.md Control-plane endpoint reference + build/run + smoke test.
agent/README.md Agent endpoints, config, profiles (owned by CP), engine install lifecycle, restart behavior.
CLAUDE.md Architecture, components, design decisions, project relationships.
docs/test-cases-todo.md Test backlog (what's automated vs. still manual).
docs/llama-cpp-as-managed-runtime.md The "Witchgrid as the llama.cpp manager" long-game (vendoring, canary/rollback).

Layout

Dir What's there
cp/ Control-plane daemon: registry, placement, routing, profiles, catalog, dashboard, GGUF introspection.
agent/ Per-node agent: hardware probe, capability introspection, engine supervisor + watchdog, install lifecycle.
dashboard/ HTMX + Alpine + Pico assets, embedded into cp via scripts/embed_assets.sh.
scripts/ Build helpers (embed_assets.sh) + the CP systemd unit template.
tests/ Unit suite (hemlock tests/run.hml) + integration harness (tests/integration/).
docs/ Design docs, ops guide, roadmaps, the Hemlock language-eval archive.
legacy-python/ Frozen v0.3 Python prototype. Reference only.

Tests

hemlock tests/run.hml                         # unit: intent / capabilities / GGUF parser
bash tests/integration/run_integration.sh     # integration: real CP+agent, fake backend

CI (.github/workflows/build.yml) builds both platforms, runs the unit + integration suites, and on a v* tag attaches per-platform binaries + SHA256SUMS to the release. See tests/README.md.

Roadmap to 1.0

Capability-wise it's there; 1.0 is about coherence, deployability, and test breadth:

  • Single source of truth for the version (version.hml), surfaced in /healthz + /metrics (witchgrid_build_info) + the dashboard header.
  • CP bind configurable — WITCHGRID_CP_HOST / WITCHGRID_CP_PORT.
  • Deploy/ops documented — see docs/operations.md.
  • Test breadth — the integration harness now covers port-contention, stale-node drop, spawn-failure, bad-model-path, multi-node registration, the CPU/GPU device override, and the live/health endpoints (26 assertions).
  • Auth posture finalized — LAN-first shared-secret model, documented in docs/security.md (what's public vs protected, reverse-proxy guidance for exposure). Deferred to post-1.0: a dashboard login (slice 2) + per-consumer /v1/llama tokens (slice 3) + native TLS.
  • Graceful shutdown — on SIGINT/TERM/HUP the CP checkpoints the WAL (TRUNCATE) and exits cleanly. Live as of the Hemlock 2.6.2 pin, which carries the async-runtime signal fix (hemlang/hemlock#587). Single-host multi-node registration is tested; a cross-node 2-agent test (mocked nvidia-smi) is the remaining test nicety.

Related projects

  • Merlina — the lab's fine-tuning UI. Merlina trains, Witchgrid serves.
  • flammen.ai — Witchgrid's first consumer (chat + image + caption workers route through it).

Built in Hemlock. The language-evaluation writeup that informed the rewrite lives in docs/hemlock-feedback-archive.md.

About

No description, website, or topics provided.

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages