A terminal coding agent that runs any model on OpenRouter with your own key.
One key, 400+ models, per-turn billed cost.
autopilot sends every model call to OpenRouter using your own OpenRouter API key — bring-your-own-key, nothing proxied, nothing marked up. On first run you paste the key once (it's checked against OpenRouter before it's saved) and pick a model. That's the whole setup.
- Any model, one key. The model picker is OpenRouter's live catalog — Claude, GPT, Gemini, Grok, Qwen, GLM, Kimi, DeepSeek, free models and more — opening on the best & latest by benchmark score, with fuzzy search over everything. Switch any time with
/model. - Real cost, per turn. OpenRouter reports what each generation actually cost; the status bar shows it (a local estimate marked
≈$only until the billed number arrives)./costtotals it by session, day, month and all time. - Reliable tool calling. Requests only go to upstream providers that support every parameter autopilot sends (tools above all), and OpenRouter routes tool-calling traffic to the providers with the best tool-call success rates. The status bar shows which provider served the last turn.
- Prompt caching that stays warm. Each session pins to one upstream provider (OpenRouter sticky routing), so the long, stable prompt prefix keeps hitting the provider's cache across turns.
- Up to 1M+ context (model-dependent) — Read entire modules, large configs, and full stack traces without the model losing track.
- Image understanding — Drop image paths (PNG, JPG, WebP, GIF, BMP up to 5 MB) into any prompt. Great for UI reviews, diagrams, and screenshots.
- Plan / Edit / Auto modes —
planis a whitelist-only research mode: only read-only tools (read, glob, grep, web search, GitHub read-only, browser fetch) are allowed. Writes, edits, mutating bash, MCP tools, and LSP renames are all blocked.edit(default) prompts per mutating call.autoapproves everything for trusted tasks. - Subagents — By default, the coordinator assesses substantial tasks for independent research and may delegate without special phrasing. Worker calls still go through permission checks and need a configured backend. See the subagents guide for examples, controls, limits, and setup.
- Long-running unattended runs — The headless server can launch detached jobs, sleep between check-ins, and resume from saved checkpoints. Overall runtime, token, cost, and action budgets are optional; explicit budget exhaustion pauses for input. See the long-running runs guide.
- Windows support — OS-aware shell auto-detects
cmd.exe/ PowerShell on Windows,bashon Unix. Thebashtool works out of the box on all platforms. - Message queuing — Submit multiple messages while the agent is busy; they queue and auto-drain. Escape interrupts the current turn but preserves the queue.
- Smart permission modal — Denying a tool opens inline feedback so you can tell the agent what to do instead. Keyboard-native navigation (
↑/↓,j/k,Alt+1/2/3). - Loop guardrails — Agent hard-stops when all tools in a turn are blocked, preventing infinite token-burning cycles.
- Persistent all-time cost history — Append-only
history.jsonltracks daily usage forever, so/costshows true all-time and monthly totals that survive across sessions and version updates. - LSP + MCP — Semantic code intelligence (hover, go-to-definition, references, diagnostics) via Language Server Protocol. Extend with external tools via Model Context Protocol.
- Local structured memory — SQLite + embeddings cross-session memory. The agent recalls facts, instructions, and preferences across sessions via
remember,recall, andforgettools. - Web search, GitHub, and headless browser — Research the web, read GitHub repos, and fetch JavaScript-rendered pages without leaving your terminal.
- OpenRouter as the model provider — Every OpenRouter model, your own key, billed cost per turn. Replaces Cloudflare Workers AI / AI Gateway (see Upgrading from Cloudflare).
- OS-aware shell with Windows support — Auto-detects
cmd.exe, PowerShell, or bash based on platform. Override withKIMIFLARE_SHELLor/shell. - Smart permission modal with inline feedback — Deny a tool and immediately tell the agent what to do instead. Keyboard-native navigation with
↑/↓,j/k,Alt+1/2/3. - True message queuing — Enter queues messages while the agent is busy; Escape interrupts and auto-drains the queue.
- Hard-stop loop guardrail — Stops token-burning cycles when all tools in a turn are blocked.
- Headless SDK — Programmatic
createAgentSessionAPI and JSONL-over-stdio RPC mode for building on top of autopilot.
See the full changelog at github.com/sinameraji/autopilot/releases.
npm install -g autopilot-ai
autopilotOn first run, Sign in with OpenRouter: approve autopilot in your browser and a key is created for you (or paste one from https://openrouter.ai/keys). Then pick a model. That's it.
Or run without installing:
npx autopilot-aiComing from
kimiflare? The package was renamed. Runnpm uninstall -g kimiflare && npm install -g autopilot-aiand useautopilotinstead ofkimiflare. Your settings, sessions, memory and cost history are kept (they stay in~/.config/kimiflareand~/.local/share/kimiflare), andKIMIFLARE_*environment variables still work.
Requires Node.js ≥ 20.
Three ways to provide it — the first one found wins:
- Environment:
OPENROUTER_API_KEY(orKIMIFLARE_OPENROUTER_KEY). With this set, the setup screen never appears — the way to run kimiflare headless (CI, a VM, a container). - Sign in / setup screen / CLI: "Sign in with OpenRouter" on first run, or
autopilot auth openrouter— approve in the browser and a key is created and saved (OpenRouter's OAuth PKCE flow). Over SSH or in a container it shows a link to open on any device and asks for the code OpenRouter displays (force this with--code).--pasteprompts for an existing key instead, andautopilot auth openrouter <key>saves one directly. Inside the TUI,/key set <key>replaces it. - Config file:
"openrouterApiKey": "sk-or-…"in~/.config/kimiflare/config.json(created with mode 600).
/key shows which key is in use, what it has spent and how much credit is left. Model calls — including memory embeddings and small internal side-calls (summaries, memory extraction) — are all billed to this key.
A key with no credits can still use OpenRouter's free models (the "Free" section of /model, ids ending in :free), subject to OpenRouter's daily request cap.
Earlier kimiflare versions ran on Cloudflare Workers AI / AI Gateway. On the first launch after upgrading you'll be asked for an OpenRouter key once; after that:
- Your settings (theme, MCP/LSP servers, hooks, memory, sessions, cost history) carry over.
- Model ids are migrated (
@cf/moonshotai/kimi-k2.6→moonshotai/kimi-k2.6, and so on), and the retired Cloudflare fields (OAuth login, gateway, Unified Billing, provider keys) are removed fromconfig.json. - Memory keeps working unchanged: embeddings use the same
bge-base-en-v1.5model, now via OpenRouter. /gateway,autopilot auth cloudflare,--cloudand the Cloud-only commands are gone./multi-agent(Commute) is retired; subagents now run locally through Hotcell.
For substantial tasks, the agent can hand independent investigations to subagents: separate Autopilot instances that each explore the repository in their own local Hotcell sandbox and report back cited findings. Several subagents run in parallel behind a single approval while the main agent keeps the conversation and does the edits. You don't need to ask for them; say "don't use subagents" to opt out, or "use subagents" to insist.
Setup: install and start Hotcell, then add your OpenRouter key to its host key store (hotcell keys add openrouter). The real key never enters a sandbox. Subagents are read-only (no edits, shell, PRs, or MCP) and see the committed, pushed state of your branch. See docs/subagents.md for limits and troubleshooting.
While subagents run they are listed above the prompt (Ink) or in the activity panel (/agents, Camouflage). /subagents cancel <n> stops one without interrupting the turn.
You pick the model on first run, and can switch any time with /model. The picker opens on the best & latest models — recent, tool-capable models ranked by the agentic and coding benchmark scores OpenRouter publishes (Artificial Analysis), at most two per lab — so it stays current as new models ship, with no list maintained in code. Start typing to fuzzy-search all of them (sonet finds Claude Sonnet, gpt 6 the GPT-6 family). Or set one directly:
/model moonshotai/kimi-k3 # 1M context
/model anthropic/claude-sonnet-5
/model deepseek/deepseek-v4-flash
autopilot -m moonshotai/kimi-k2.7-code -p "..."Any OpenRouter model id works, including variants like :free and :nitro. Models OpenRouter lists without tool calling are hidden from the picker — a coding agent needs tools.
Provider routing (optional). OpenRouter serves most models through several upstream providers. autopilot always requires providers that support every parameter it sends, and otherwise leaves routing to OpenRouter (price-weighted, uptime-aware, tool-calling quality first, sticky per session for caching). To steer it, add an openrouterProvider object to the config file — any of OpenRouter's provider preferences:
{ "openrouterProvider": { "ignore": ["SomeProvider"], "data_collection": "deny" } }Avoid order / sort unless you need them: they turn off sticky routing, which is what keeps the prompt cache warm.
Point every model call at your own OpenAI-compatible endpoint instead of OpenRouter — useful when a host application (CI, an agents platform, a container) fronts model access with its own broker and doesn't want to hand autopilot a raw provider key:
export KIMIFLARE_BASE_URL="https://your-broker.example.com/v1" # /chat/completions is appended
export KIMIFLARE_API_KEY="<bearer for that endpoint>" # optional; header omitted if unset
autopilot -p "..." # or --mode rpc — no OpenRouter key neededThe same pair can be persisted in ~/.config/kimiflare/config.json as baseUrl / apiKey
(env vars win over the file, field by field). When a base URL is configured:
- Chat requests go to
<baseUrl>/chat/completionsand memory embeddings to<baseUrl>/embeddings, withAuthorization: Bearer $KIMIFLARE_API_KEY(noAuthorizationheader at all when the key is unset — e.g. a local llama.cpp/Ollama server). - Model ids pass through in the request body unchanged; your endpoint owns provider dispatch.
- Cost is the local estimate only — your endpoint does its own metering.
Requesty is another OpenAI-compatible gateway you can use with your own key instead of OpenRouter. It is opt-in: it is only used when no OpenRouter key and no custom endpoint are configured, so an existing setup never changes.
export REQUESTY_API_KEY="<your key>" # from https://app.requesty.ai/api-keys
autopilot # or: autopilot auth requesty (checks and saves the key)
autopilot -p "..." -m openai/gpt-4o-mini- The key can also live in
~/.config/kimiflare/config.jsonasrequestyApiKey(the env var wins). REQUESTY_BASE_URLpicks a region, e.g.https://router.eu.requesty.ai/v1for the EU. Only https URLs on Requesty's hosts (router.requesty.ai,router.eu.requesty.ai,router.us.requesty.ai,router.ap.requesty.ai) are accepted, so the key is never sent anywhere else./modellists Requesty's managed policies first (short ids such askimi-k2.6orclaude-sonnet-4-5), then its fullvendor/modelcatalog. Both kinds of id work.- Defaults on Requesty:
kimi-k2.6for the main model,deepseek-v4-flashfor side calls, andopenai/text-embedding-3-smallfor memory embeddings. - Cost per turn is the one Requesty reports inline.
/keyshows whether the key is valid; OpenRouter only features (browser sign in,/key set, the billed cost lookup) are not available.
autopilot -p "summarize PLAN.md" # stream answer to stdout
autopilot -p "..." --dangerously-allow-all # auto-approve mutating tools (for scripts)
autopilot -p "..." --reasoning # include chain-of-thought in stderrUse autopilot programmatically from your own application — no TUI required.
import { createAgentSession } from "autopilot-ai/sdk";
const { session } = await createAgentSession({
cwd: "/path/to/project",
config: {
openrouterApiKey: process.env.OPENROUTER_API_KEY,
model: "moonshotai/kimi-k2.7-code",
},
});
// Stream every event: text deltas, tool calls, tasks, usage
session.subscribe((event) => {
console.log(event.type, event);
});
// Send a prompt
await session.prompt("Refactor auth to JWT + Redis");
// Mid-flight correction while the agent is still running
await session.steer("Use Redis instead of in-memory store");
// After the turn finishes
await session.followUp("Also add unit tests");
// Clean up
session.dispose();Key features:
subscribe()— receive typed events (text_delta,tool_call,tool_result,task_update,usage,warning,error,done, etc.)prompt()/steer()/followUp()— full conversation lifecyclepause()/resume()— graceful preemptiongetStatus()/getUsage()— inspect session state- Custom
permissionHandler— decide programmatically whether to allow mutating tools - Optional
memoryEnabled,lspEnabled,costAttributionflags
The SDK needs an OpenRouter API key — or a custom gateway endpoint
(baseUrl / apiKey in config, or KIMIFLARE_BASE_URL / KIMIFLARE_API_KEY), in which case no
OpenRouter key is required. Resolved in this priority order:
- Explicit
configobject (openrouterApiKey) — recommended for apps - Environment variables:
OPENROUTER_API_KEY/KIMIFLARE_OPENROUTER_KEY - Config file:
~/.config/kimiflare/config.json
A Requesty key (requestyApiKey in config, or REQUESTY_API_KEY) also works
when no OpenRouter key is set; pass a Requesty model id such as openai/gpt-4o-mini as model.
Pass provider in createAgentSession options to set OpenRouter provider-routing preferences for that session.
For Electron / desktop apps, we recommend storing the key in the OS keychain (e.g. Electron safeStorage or keytar) and passing it explicitly:
import { createAgentSession } from "autopilot-ai/sdk";
const openrouterApiKey = await keytar.getPassword("autopilot", "openrouter");
const { session } = await createAgentSession({
cwd: projectPath,
config: { openrouterApiKey },
});If you need process isolation or a non-Node consumer, run autopilot in JSONL-over-stdio RPC mode:
node bin/autopilot.mjs --mode rpcGive the subprocess OPENROUTER_API_KEY — or a custom gateway endpoint
(KIMIFLARE_BASE_URL + KIMIFLARE_API_KEY), ideal for host apps that broker model access themselves.
import { spawn } from "node:child_process";
const proc = spawn("npx", ["autopilot-ai", "--mode", "rpc"], {
cwd: projectPath,
stdio: ["pipe", "pipe", "pipe"],
});
// Read events
proc.stdout.on("data", (chunk) => {
for (const line of chunk.toString().split("\n")) {
if (!line.trim()) continue;
const event = JSON.parse(line);
console.log(event.type, event);
}
});
// Send commands
proc.stdin.write(JSON.stringify({ type: "new_session" }) + "\n");
proc.stdin.write(JSON.stringify({ type: "prompt", message: "Hello" }) + "\n");
// Resolve a permission request
proc.stdin.write(
JSON.stringify({ type: "resolve_permission", requestId: "req_0", decision: "allow" }) + "\n"
);
// Resume a previous session after a process restart (the `new_session`
// response echoes back the sessionId to store for later)
proc.stdin.write(JSON.stringify({ type: "new_session", sessionId: "sdk-session-…" }) + "\n");autopilot
› fix the layout bug in this screenshot docs/bug.png
› convert this mockup design.png to Tailwind HTML| Command | Effect |
|---|---|
/mode edit|plan|auto |
Switch permission mode |
/shell auto|bash|cmd|powershell |
Show or set the shell for the bash tool |
/thinking low|medium|high |
Reasoning effort (persists) |
/theme |
Interactive theme picker (Ctrl+T) |
/resume |
Pick a past conversation to restore |
/compact |
Summarize older turns to free context |
/init |
Scan repo and write KIMI.md project context |
/memory |
Show memory stats and search |
/mcp list / /mcp reload |
Manage MCP servers |
/reasoning |
Toggle chain-of-thought display |
/model |
Pick a model from OpenRouter's catalog (or /model <id>, /model list [filter]) |
/jev |
Ask Jev for a typed, scoped one-shot decision |
/subagents |
List running subagents; /subagents cancel <n>|all stops them; /subagents off|suggest|auto sets the delegation policy (default: auto) |
/now |
Run your latest queued message now instead of after the current turn (also Ctrl+G in Ink) |
/key |
Show your OpenRouter key's spend and credit (/key set <key>, /key clear) |
/cost |
Show OpenRouter-confirmed cost by session, day, month and all time |
/update |
Check for updates |
/help |
List all commands |
/subagents auto is the default: on substantial tasks the agent is asked to look for independent parts and delegate them, and it decides whether delegation is worth it. suggest only hints at delegation for likely parallel work, and off disables harness guidance (you can still ask for subagents explicitly). Explicit requests to use subagents or to work without them always win. Set the policy with /subagents, KIMIFLARE_SUBAGENT_POLICY, or subagentPolicy in config.
The harness never starts a subagent by itself: the agent calls the subagent tool, which always needs normal tool permission (one prompt covers a parallel batch) and obeys provider, read-only, spend, concurrency, and timeout limits. In Code Mode the subagent tool stays a direct tool next to execute_code. Jev is consulted only in suggest mode for substantial but ambiguous requests, receives at most 1,200 characters of redacted task text, and has a 4-second timeout.
Type ! followed by a command (for example ! gh auth login) to run it in your own terminal, inside the session. Logins, password prompts and interactive programs work because the command gets the real TTY. The output is added to the conversation so the agent can see it. The agent's own bash tool has no terminal, so autopilot suggests ! <command> when something needs you: a login, a sudo password, an SSH passphrase, or a secret you'd rather not paste into chat. On macOS and Linux the output is captured with script(1); elsewhere the command still runs but its output isn't captured.
| Shortcut | Action |
|---|---|
Ctrl+C / Esc |
Interrupt current turn when busy; exit when idle |
Ctrl+R |
Toggle reasoning display |
Ctrl+O |
Toggle verbose tool output |
Ctrl+T |
Open theme picker |
Shift+Tab |
Cycle mode (edit → plan → auto) |
↑ / ↓ |
Walk prompt history |
Ctrl+G |
While the agent works: run your latest queued message now instead of after the turn |
A message you send mid-turn is triaged instead of always waiting in line:
- interrupt ("stop, that's the wrong file"): the turn stops at the next safe point and your message runs next. Running subagents are never cancelled by an interrupt; while they run, the message is delivered as soon as they finish.
- steer ("also handle the empty case"): folded into the current task at the agent's next step.
- queue ("after this, update the changelog", or anything unclear): runs as its own turn afterwards. The queue shows
Ctrl+G run this nowso you can promote it. - aside ("how's it going?"): answered from the current state without disturbing the turn.
Clear phrasing is decided instantly; ambiguous messages get a quick model check (OpenRouter only), and anything uncertain is queued. Use /subagents cancel <n> to stop a single subagent.
autopilot writes structured JSON logs of agent-side activity (tool calls,
permission decisions, MCP/LSP lifecycle, session events, errors) to
~/.config/kimiflare/logs/<date>.jsonl, one file per day, with 7-day
retention pruned automatically at startup.
The logs deliberately exclude prompts and completions. To capture full
request/response payloads locally for debugging, run with
KIMIFLARE_DUMP_LLM=1; each generation is also visible in your
OpenRouter activity log.
autopilot logs path # today's file
autopilot logs dir # log directory
autopilot logs prune # delete files older than 7 days
# Tail this session's activity, formatted:
tail -f $(autopilot logs path) | jq
# Find the slowest tool calls in the last day:
jq -r 'select(.event == "tool:end") | "\(.data.duration_ms)\t\(.data.tool)"' \
$(autopilot logs path) | sort -rn | headDisable the file sink entirely with KIMIFLARE_LOG_SINK=off. The
separate KIMIFLARE_LOG_LEVEL env var (default off) controls stderr
output — independent of the file sink.
If you set KIMIFLARE_OTEL_ENDPOINT, autopilot also ships each log
entry to that endpoint over OTLP/HTTP
so it lands in Datadog, Honeycomb, Grafana Loki, an internal collector,
or any other backend that speaks OTel. Batched every 5 s (or every
100 entries, whichever first) and best-effort — never blocks the agent
loop.
# Full path:
export KIMIFLARE_OTEL_ENDPOINT="https://otel.example.com/v1/logs"
# Or just the base URL (we auto-append /v1/logs):
export KIMIFLARE_OTEL_ENDPOINT="https://otel.example.com"
# Optional headers (comma-separated key=value pairs) — e.g. for auth:
export KIMIFLARE_OTEL_HEADERS="Authorization=Bearer xyz,X-Tenant=acme"Each log entry maps to one OTel LogRecord. Correlation IDs
(session_id, turn_id, request_id) become record attributes,
data.* fields are flattened to attributes with type-preserving
encoding, and a service.name=kimiflare + service.version pair sits
on the resource.
autopilot can fire shell commands at five points in an agent turn,
configured per-project (.kimiflare/settings.json) or globally
(~/.config/kimiflare/settings.json):
| Event | Fires when | Veto? |
|---|---|---|
PreToolUse |
A tool call is about to run | Yes |
PostToolUse |
A tool call just finished | No |
UserPromptSubmit |
You hit Enter on a prompt | Yes |
Stop |
A turn ended cleanly | No |
PreCompact |
Auto-compaction is about to run | No |
Hooks receive the event payload as JSON on stdin and as
KIMIFLARE_HOOK_* env vars (for shell-one-liner ergonomics).
Non-zero exit on a veto event cancels the underlying action and
surfaces the hook's stdout as the rejection reason.
/hooks # list configured hooks
/hooks recommended # list starter hooks shipped with autopilot
/hooks enable stop-bell # enable one (writes to .kimiflare/settings.json)
/hooks enable stop-bell global # ...or the global file
/hooks disable stop-bell
/hooks path # print settings.json paths
/hooks reload # re-read settings.json after a manual edit
The recommended catalog includes terminal bells / macOS notifications
on Stop, secret-file guards on PreToolUse (e.g. block edits to
*.env), auto-format-with-prettier on PostToolUse, and a tool-call
audit log. All ship disabled — /hooks recommended lists them.
{
"hooks": {
"PreToolUse": [
{
"id": "no-secrets",
"matcher": "^(edit|write)$",
"command": "case \"$KIMIFLARE_HOOK_PATH\" in *.env|*.pem) echo 'blocked'; exit 1;; esac"
}
],
"PostToolUse": [
{
"id": "format-ts",
"matcher": "^(edit|write)$",
"command": "npx --no-install prettier --write \"$KIMIFLARE_HOOK_PATH\" >/dev/null 2>&1 || true"
}
],
"Stop": [
{ "id": "bell", "command": "printf '\\a'" }
]
}
}Per-hook fields:
command(required) — the shell command.matcher(optional) — anchored regex matched against the tool name forPreToolUse/PostToolUse. Ignored for other events.id(optional) — stable handle for/hooks enable|disable. Auto-derived fromevent + commandwhen omitted.enabled(defaulttrue) — setfalseto keep a hook in config but skip it.timeoutMs(default30000) — hard kill if the hook hangs.description(optional) — shown by/hooks list.
Hooks are always-on infrastructure: they fire whether the TUI is open
or autopilot is running in --print mode. They also fire for tool
calls generated from inside the Code Mode sandbox (heavy-tier turns),
because hook firing lives on the ToolExecutor itself — every call
path uses the same plumbing.
When intent classification has assigned a tier, hook payloads include
it as tier: "light" | "medium" | "heavy" (on UserPromptSubmit,
PreToolUse, PostToolUse) and as $KIMIFLARE_HOOK_TIER. Useful for
"skip auto-format on light turns" or "audit every heavy-turn write."
SDK consumers opt in to hooks with enableHooks: true on
createAgentSession. Default is off because the SDK is a primitive,
not the TUI.
git clone https://github.com/sinameraji/autopilot
cd autopilot
npm install
npm run build
npm linkScripts:
npm run build— bundle with tsupnpm run dev— run via tsxnpm run typecheck—tsc --noEmitnpm test— run tests
- Fork the repository
- Create a branch:
git checkout -b feat/your-feature - Make your changes
- Run
npm run typecheckandnpm run build - Commit with Conventional Commits
- Open a Pull Request
Built by Sina Meraji and contributors · MIT License
