English · 简体中文
All Anthropic OAuth stripped. All telemetry stripped. All injected security-prompt guardrails removed. All experimental features unlocked. One binary, zero callbacks home.
curl -fsSL https://raw.githubusercontent.com/Flybicy/CC-lite/main/install.sh | bashOne command installs all dependencies first (git, Bun >= 1.3.11, ripgrep), then clones the source, builds, provisions the local semantic embedding runtime and pre-downloads the model (~23 MB), and installs
cclite,cclite-bypass(bypass permission mode), andcclite-verify-embeddings(re-checks the semantic model). See API Configuration for API setup.
curl -fsSL https://raw.githubusercontent.com/Flybicy/CC-lite/main/install_dev.sh | bashInstalls from the
devbranch ascclite(pluscclite-bypassandcclite-verify-embeddings). Same command names as the stable installer — only the source branch differs.
This is a clean, buildable fork of Anthropic's Claude Code CLI -- the terminal-native AI coding agent. The upstream source became publicly available on March 31, 2026 through a source map exposure in the npm distribution.
This fork applies six categories of changes on top of that snapshot:
Eliminates all tracking and remote-control mechanisms present in the original Claude Code:
- No telemetry -- No unnecessary data is transmitted to Anthropic servers
- No analytics -- No usage tracking or event logging
- No fingerprinting -- No user or environment identification
- No auto-updates -- No remote version control or forced updates
Unlike the upstream Claude Code, CC-lite has no OAuth login, no claude.ai remote sessions, and no cloud provider integration:
- No
/logincommand -- authenticating with claude.ai OAuth is removed - No remote CCR sessions -- all bridge/remote session code is stripped
- No GrowthBook server-side feature flag dependency
- No auto-update infrastructure
- No settings sync to/from cloud
All authentication is done via API keys (see API Configuration).
Added an API shim layer (src/services/api/openaiShim.ts) that transparently translates between Anthropic message format and OpenAI-compatible APIs. It supports both Chat Completions and the newer Responses API, so all CC-lite tools (bash, file read/write, grep, glob, agents, MCP, etc.) keep working while you swap in a different backend LLM.
The OpenAI transports honor configured request timeouts and keep Codex-style originator, OpenAI-Beta, and session_id headers scoped to the Responses transport. The /goal completion judge receives the active model explicitly, so model switching does not leave goal checks using stale or invalid runtime state.
The current main branch includes request-timeout handling for OpenAI transports, Responses-only Codex headers, and explicit active-model forwarding for /goal. The maintenance baseline is 273 pass / 31 files with bun test; build the distributable bundle with bun run build:bundle:cclite.
Added an optional override for the built-in WebSearch tool so it can query your own SearXNG instance instead of relying on provider-side web search.
- Configured with one env var:
CLAUDE_CODE_SEARXNG_BASE_URL - Keeps the existing
WebSearchtool contract and UI intact - Preserves
allowed_domains/blocked_domainssemantics with local filtering - Falls back to the default provider behavior when the env var is unset
Anthropic injects system-level instructions into every conversation that constrain Claude's behavior beyond what the model itself enforces. These include:
- Hardcoded refusal patterns for certain categories of prompts
- Injected "cyber risk" instruction blocks
- Managed-settings security overlays pushed from Anthropic's servers
This build strips those injections. The model's own safety training still applies -- this just removes the extra layer of prompt-level restrictions that the CLI wraps around it.
Claude Code ships with dozens of feature flags gated behind bun:bundle compile-time switches. Most are disabled in the public npm release. This build unlocks every flag that both compiles cleanly and actually works without claude.ai OAuth, including:
| Feature | What it does |
|---|---|
ULTRATHINK |
Deep thinking mode -- type "ultrathink" to boost reasoning effort |
VOICE_MODE |
Push-to-talk voice input and dictation |
AGENT_TRIGGERS |
Local cron/trigger tools for background automation |
TOKEN_BUDGET |
Token budget tracking and usage warnings |
BUILTIN_EXPLORE_PLAN_AGENTS |
Built-in explore/plan agent presets |
VERIFICATION_AGENT |
Verification agent for task validation |
BASH_CLASSIFIER |
Classifier-assisted bash permission decisions |
EXTRACT_MEMORIES |
Post-query automatic memory extraction |
HISTORY_PICKER |
Interactive prompt history picker |
MESSAGE_ACTIONS |
Message action entrypoints in the UI |
QUICK_SEARCH |
Prompt quick-search |
SHOT_STATS |
Shot-distribution stats |
COMPACTION_REMINDERS |
Smart reminders around context compaction |
CACHED_MICROCOMPACT |
Cached microcompact state through query flows |
Flags that hard-depend on stripped claude.ai infrastructure --
ULTRAPLAN(remote CCR planning, needs/login) andBRIDGE_MODE(Remote Control bridge, gated on claude.ai subscription) -- are excluded on purpose, along with the other OAuth-only flags (CCR_*,AGENT_TRIGGERS_REMOTE). Enabling them would only ship dead code. See FEATURES.md.
See FEATURES.md for the full audit of all 88 flags and their status.
This repo ships ready for GitHub. To publish your own copy:
# 1) Create an empty repo on GitHub (e.g. github.com/you/cclite), then:
cd cclite
git remote set-url origin https://github.com/you/cclite.git
git push -u origin mainTo pull updates and rebuild:
git pull --ff-only # fast-forward only; never auto-merge
bun install # sync deps if package.json/lock changed
bun run build:dev:cclite # rebuild cclite-cli-devIf you have local changes and want to re-sync with upstream cleanly:
git stash # park your edits
git pull --ff-only
git stash pop # reapplycurl -fsSL https://raw.githubusercontent.com/Flybicy/CC-lite/main/install.sh | bashOne command does everything, in this order:
- Dependencies first -- git, Bun >= 1.3.11, and ripgrep, via your system package manager (apt / dnf / yum / pacman / zypper / apk / brew; on Windows Git Bash: scoop / choco / winget or a direct release download).
- Clone + build -- clones the repo source, runs
bun install, and builds the JS bundlecclite.jswith all experimental features enabled. - Semantic model, in one step -- installs the embedding runtime
(Transformers.js + ONNX Runtime) into
~/.local/lib/cclite, trims it to the host platform (~130 MB), pre-downloads the ~23 MB model, and runs a real similarity smoke test so you see semantic search verified during install. No manual setup, no first-run download stall. - Launchers + PATH -- installs
cclite,cclite-bypass, andcclite-verify-embeddingsinto~/.local/bin, and adds that directory to your shell profiles (~/.bashrc/~/.zshrc/~/.profile) so the command works immediately in new terminals.
Why a bundle instead of one compiled binary?
bun build --compileplaces ONNX Runtime's native.nodelibrary inside the executable's virtual filesystem, wheredlopen()cannot load it -- so a single-file binary can only ever use the approximate fallback, silently. Shippingcclite.jsplus a small runtime directory is what makes the real semantic model work. Startup cost is negligible (~0.5s).
CC-lite builds and runs natively on Windows (no WSL needed). Use the
PowerShell installer, which follows the same order: all dependencies first
(git, Bun, and ripgrep -- via winget/choco/scoop), then clone + bun install
- build the bundle, then provision the semantic embedding runtime and
pre-download/verify the model, then install
cclite.cmd,cclite-bypass.cmd, andcclite-verify-embeddings.cmd:
irm https://raw.githubusercontent.com/Flybicy/CC-lite/main/install.ps1 | iexOr, if you already have Bun + ripgrep, build from source as on any other OS
(see Build). ripgrep (rg) must be on your PATH -- CC-lite
does NOT bundle it; install it with scoop install ripgrep,
winget install BurntSushi.ripgrep.MSVC, or choco install ripgrep.
Note: ripgrep is not bundled by CC-lite; it's a system dependency on all platforms (brew/apt install ripgrep on macOS/Linux).
- Bun >= 1.3.11
- macOS, Linux, or Windows (native build; WSL optional but not required)
- An API key (Anthropic Messages or OpenAI-compatible APIs)
# Install Bun if you don't have it
curl -fsSL https://bun.sh/install | bash# Clone the repo
git clone https://github.com/Flybicy/CC-lite.git
cd cclite
# Install dependencies
bun install
# Standard build -- produces ./cclite-cli
bun run build
# Dev build -- dev version stamp, experimental GrowthBook key
bun run build:dev
# Dev build with ALL experimental features enabled -- produces ./cclite-cli-dev
bun run build:dev:full
# Compiled build (alternative output path) -- produces ./dist/cclite-cli
bun run compile
# Bundle build -- produces ./cclite.js, the variant the installers ship.
# Keeps the embedding stack external so the REAL semantic model works.
bun run build:bundle:cclite
# Verify the local semantic model end-to-end (downloads it on first run)
bun run verify:embeddings| Command | Output | Features | Notes |
|---|---|---|---|
bun run build |
./cclite-cli |
VOICE_MODE only |
Production-like binary |
bun run build:dev |
./cclite-cli-dev |
VOICE_MODE only |
Dev version stamp |
bun run build:dev:full |
./cclite-cli-dev |
All working experimental flags | The full unlock build |
bun run compile |
./dist/cclite-cli |
VOICE_MODE only |
Alternative output directory |
bun run build:bundle:cclite |
./cclite.js |
cclite feature set | What the installers ship. Needs a sibling node_modules for the embedding stack; the only variant where the real semantic model runs |
The compiled single-file variants (
cclite-cli,cclite-cli-dev) cannot load ONNX Runtime's native library out of their virtual filesystem, so semantic search degrades to the approximate fallback there. Use the bundle build (orbun run dev) for true semantic search.
You can enable specific flags without the full bundle:
# Enable just ultrathink and token budget
bun run ./scripts/build.ts --feature=ULTRATHINK --feature=TOKEN_BUDGET
# Enable a specific flag on top of the dev build
bun run ./scripts/build.ts --dev --feature=MCP_RICH_OUTPUT# Run the installed command
cclite
# Run in bypass permission mode
cclite-bypass
# Re-check the local semantic embedding model
cclite-verify-embeddings
# Or the bundle you built yourself (real semantic search)
bun ./cclite.js
# Or a compiled binary (semantic search falls back to approximate)
./cclite-cli-dev
# Or run from source without building (slower startup)
bun run dev
# See [API Configuration](#api-configuration) for API setup.cclite-bypass is installed by install.sh. It exports IS_SANDBOX=1 and runs cclite with --permission-mode bypassPermissions. cclite-verify-embeddings re-runs the local semantic model check (see Verifying the semantic model).
# One-shot mode
cclite -p "what files are in this directory?"
# One-shot mode with bypass permission mode enabled
cclite-bypass -p "scan this repo and summarize risky scripts"
# Interactive REPL (default)
cclite
# Interactive REPL (bypassPermissions)
cclite-bypass
# With specific model
cclite --model claude-sonnet-4-6-20250514
# Set advisor model
cclite --advisor-model claude-sonnet-4-6-20250514CC-lite includes a built-in Advisor tool that runs a stronger reviewer model to audit your approach before you commit to implementation. The advisor checks for architecture flaws, security issues, edge cases, and correctness.
- Configured via
CLAUDE_CODE_ADVISOR_MODELor the/advisor <model>slash command - Provider-agnostic — works with any model, including OpenAI-compatible backends
- Automatically called by the executor model on the first significant action of each task
- Also available for manual review: call the
Advisortool with your question
# Enable via env var
export CLAUDE_CODE_ADVISOR_MODEL="claude-opus-4-6"
# Enable via slash command (interactive session)
/advisor claude-opus-4-6
# Disable
/advisor offThe Advisor reads the main agent's conversation history through its
ReadConversationLog tool. It now supports three search modes when using
action: "search":
| mode | what it does | when to use |
|---|---|---|
hybrid (default) |
reciprocal-rank fusion (RRF, k=60) of keyword + semantic — the local embedding model participates on every search | best for most questions |
keyword |
BM25 exact-term search | identifiers, file names, error codes, exact API names |
semantic |
embedding-vector cosine similarity only | the topic may be discussed using different words than your query |
The Advisor is actively prompted to pick a mode deliberately: keyword for exactness, semantic for paraphrase/concept recall, hybrid for broad recall. Search output always labels the embedding backend in use, so the Advisor knows which one ran and never gets silent fallback.
Semantic and hybrid modes need embeddings. CC-lite runs them fully local -- no API key, no per-token cost, your conversation text never leaves the machine. Two backends, resolved in this priority order:
- local-semantic (default, true semantics) -- a real embedding model
running in-process via
@huggingface/transformers(ONNX/WASM, L2-normalized; mean-pooled for MiniLM, CLS-pooled + query-side instruction prefix for bge models per their official retrieval recipe), labeledlocal-semantic:<model>. The installer sets this up completely: it provisions the runtime, pre-downloads the model, and verifies real inference before finishing -- nothing to configure. From source,bun installis enough. Everything after the one-time ~23 MB download works fully offline. Vectors are additionally cached on disk (JSONL per model), so unchanged messages are never re-embedded, even across restarts.- Locale-aware Chinese default: on a zh locale the default model is
Xenova/bge-small-zh-v1.5(far better Chinese retrieval than MiniLM); everything else gets MiniLM. An explicitCLAUDE_CODE_ADVISOR_LOCAL_EMBEDDING_MODELalways wins. - Download progress: the first fetch prints live percent/bytes progress.
- Startup warm-up: the interactive REPL preloads the pipeline in the
background ~2s after launch, so the first search pays zero load cost.
Note: the standalone
--compilebinary cannot load ONNX Runtime's native library from its virtual filesystem, so semantic search degrades there -- this is exactly why the installers ship thecclite.jsbundle instead.
- Locale-aware Chinese default: on a zh locale the default model is
- local-approximate (fallback) -- a deterministic, offline, free,
hashed bag-of-features vectorizer. It is NOT true semantics; it
provides fuzzy sub-word matching on top of what BM25 already does. Always
labeled
local-approximatein tool output, so the Advisor is told to lean toward keyword mode for exactness when it sees this label. Used when you opt out of the model tier (CLAUDE_CODE_ADVISOR_LOCAL_EMBEDDING=0).
Model download sizes (quantized q8, one-time, includes tokenizer/config):
| Model | Size | Notes |
|---|---|---|
Xenova/all-MiniLM-L6-v2 (default outside zh locales) |
~23 MB | fast, English-leaning |
Xenova/bge-small-zh-v1.5 (auto-selected on zh locales) |
~23 MB | recommended for Chinese, CLS pooling + query prefix applied automatically |
Xenova/multilingual-e5-small |
~120 MB | strongest multilingual recall |
Pick a model with CLAUDE_CODE_ADVISOR_LOCAL_EMBEDDING_MODEL. The download
lands in the cache dir (next to the vector cache) and is reused forever.
| Variable | Description |
|---|---|
CLAUDE_CODE_ADVISOR_LOCAL_EMBEDDING_MODEL |
Transformers.js/ONNX model id for the local-semantic tier. When unset, picked by locale: Xenova/bge-small-zh-v1.5 on Chinese systems, Xenova/all-MiniLM-L6-v2 elsewhere. Downloaded once on first use, then offline. |
CLAUDE_CODE_ADVISOR_LOCAL_EMBEDDING |
Set to 0, false, or off to skip the model tier and use the approximate local fallback instead. |
CLAUDE_CODE_ADVISOR_SEMANTIC_SEARCH |
Set to 0, false, or off to disable semantic/hybrid modes entirely (mode then ignored, keyword used). Alias: CLAUDE_CODE_SEMANTIC_SEARCH. |
CLAUDE_CODE_ADVISOR_EMBEDDING_CACHE_DIR |
Override the on-disk cache directory for models and vectors (useful for tests / ephemeral CI). |
Example (Chinese conversations): on a zh locale nothing to configure --
echo $LANG # zh_CN.UTF-8 → Xenova/bge-small-zh-v1.5 picked automatically
# first semantic search downloads ~23 MB (with progress); offline after thatIf the model tier is unavailable (e.g. inside the compiled binary without the opt-out), semantic/hybrid modes report a clear error and hybrid degrades to keyword-only; keyword mode is always available regardless.
The installer already runs this check and prints the result, but you can re-run it any time:
cclite-verify-embeddings # installed alongside cclite
# or, from a source checkout:
bun run verify:embeddingsExpected output on a healthy install:
[+] Model ready in 1.2s - backend: local-semantic:Xenova/all-MiniLM-L6-v2
dimensions: 384
similarity related: 0.659 unrelated: 0.089
[+] Semantic search verified: real embeddings are working
The check embeds a query plus one related and one unrelated sentence and
asserts the related score clearly outranks the unrelated one -- the
approximate fallback cannot pass it. A second source of truth: the Advisor's
search output always prints [embedding backend: ...], which reads
local-semantic:<model> when the real model is active and
local-approximate when it is not.
The Advisor's conversation log also keeps a per-project long-term memory:
every snapshot is persisted (de-duplicated by content fingerprint) to a small
JSONL archive keyed by the project directory. After a restart, or in a
brand-new session in the same project, the Advisor can still search and read
what was discussed before -- prior entries appear in index tagged
(prior session) with high ids (>= 1000000), and work with every action
(search, read, around, both keyword and semantic modes).
| Variable | Description |
|---|---|
CLAUDE_CODE_ADVISOR_PROJECT_MEMORY |
Set to 0, false, or off to disable the per-project archive entirely. |
CLAUDE_CODE_ADVISOR_PROJECT_MEMORY_DIR |
Override the archive directory (default: a per-project hash dir under the OS cache dir). |
The archive is FIFO-bounded (newest 4000 entries), never duplicates unchanged messages across runs/resumes, and all persistence is best-effort -- a failing disk never breaks the Advisor.
The /resume (and --resume) session picker scores candidate sessions with
the same local embedding model in addition to its keyword filter: when a
search dead-ends on titles/branches/tags, topically related sessions still
appear at the end of the list (up to 12 extras, cosine >= 0.3). Fully
offline, zero API cost, best-effort -- when embeddings are unavailable the
picker behaves exactly as before. Session vectors share the disk cache, so
repeat searches are nearly free.
- Copy on select: finishing a mouse selection (drag or multi-click) copies
it automatically -- OSC 52 / tmux / pbcopy / wl-copy fallbacks. Disable via
copyOnSelectin/config. - In-app clipboard floor: every copy is also mirrored into CC-lite's internal clipboard, so copying always works even when the terminal lacks OSC 52 support or wl-copy/xclip.
- ctrl+y pastes into the prompt: inserts the most recent copy at the
cursor (readline yank; passes through over SSH/tmux everywhere). Long text
folds into a paste pill. Rebind via
chat:yankinkeybindings.json.
scripts/
build.ts # Build script with feature flag system
src/
entrypoints/cli.tsx # CLI entrypoint
commands.ts # Command registry (slash commands)
tools.ts # Tool registry (agent tools)
QueryEngine.ts # LLM query engine
screens/REPL.tsx # Main interactive UI
commands/ # /slash command implementations
tools/ # Agent tool implementations (Bash, Read, Edit, etc.)
components/ # Ink/React terminal UI components
hooks/ # React hooks
services/ # API client, MCP, analytics
state/ # App state store
utils/ # Utilities
skills/ # Skill system
plugins/ # Plugin system
bridge/ # IDE bridge
voice/ # Voice input
tasks/ # Background task management
| Runtime | Bun |
| Language | TypeScript |
| Terminal UI | React + Ink |
| CLI parsing | Commander.js |
| Schema validation | Zod v4 |
| Code search | ripgrep (required on PATH) |
| Protocols | MCP, LSP |
| API | Anthropic Messages API / OpenAI-compatible APIs |
CC-lite supports both Anthropic Messages API (natively) and OpenAI-compatible APIs. The shim can use either Chat Completions or the newer Responses API depending on the provider and model you configure.
The easiest way to configure providers is the local WebUI:
ccliteweb # opens http://127.0.0.1:1511
# (alias: `cclite config` and `cclite web` still work)The page lets you:
-
register and keep any number of providers (OpenAI-compatible or Anthropic-compatible; local servers like Ollama/LM Studio work too — API key optional there),
-
pull each provider's model list with one click (
GET /models), -
bind the pro / plus / se codenames to a provider + model each:
codename failover position suggestion promain slot — serves the conversation by default your strongest model plusfirst failover when pro fails a mid-priced model sesecond failover when plus also fails a cheap or local model The codebase always calls models by codename (
/model pro,--model se), so what a codename points at is decided entirely in the WebUI — mix vendors freely.
Configuration is stored at ~/.claude/providers.json (plain JSON, 0600
permissions). The server binds 127.0.0.1 only (with Host/Origin checks that
reject DNS rebinding) and stops when the ccliteweb command exits. Saves
are hot-reloaded: a running CLI picks up edits on its next request, no
restart. Port precedence is --port <n> > CCLITE_CONFIG_PORT > 1511; if the
port is taken the server scans upward for a free one and prints the real URL.
--no-open skips the browser, and it is skipped automatically on headless
Linux/SSH sessions.
Tier bindings win over the env vars below; unbound tiers fall back to the env-driven behavior, so existing setups keep working unchanged.
When the bound model fails after retries (timeouts, 5xx, 429), the request
transparently falls back one tier at a time: pro → plus → se. After a
transient fallback succeeds, the chain climbs back to the original tier on
the next iteration (retry budget caps at 5 automatic downgrades per
session). Balance/quota errors (402 / "credit balance too low" style) are
sticky: the downgrade stays for the rest of the session and never climbs
back — recharge first, then switch with /model. 4xx user errors (bad
params, 401/403 auth, 404) do NOT trigger fallback; they fail fast rather
than masking the real config problem under a second vendor's error. CLI
flag --fallbackModel <id> overrides the chain when set.
Note: Unlike the upstream Claude Code, CC-lite does not support OAuth login via claude.ai. All authentication is done via API keys.
Use the official Anthropic Messages API with your Anthropic API key.
export ANTHROPIC_API_KEY="sk-ant-..."Enable OpenAI-compatible mode and configure your preferred provider:
export CLAUDE_CODE_USE_OPENAI=1| Variable | Description |
|---|---|
OPENAI_API_KEY |
API key (required for cloud APIs, optional for local models) |
OPENAI_BASE_URL |
API base URL (default: https://api.openai.com/v1) |
OPENAI_MODEL |
Model ID (default: gpt-4o) |
OPENAI_API_MODE |
Force transport selection: chat_completions or responses |
ANTHROPIC_DEFAULT_OPUS_MODEL |
Override which concrete model the opus alias resolves to |
ANTHROPIC_DEFAULT_SONNET_MODEL |
Override which concrete model the sonnet alias resolves to |
ANTHROPIC_DEFAULT_HAIKU_MODEL |
Override which concrete model the haiku alias resolves to |
CLAUDE_CODE_MAX_CONTEXT_TOKENS |
Override max context window size |
CLAUDE_CODE_SUMMARY_OUTPUT_TOKENS |
Override token limit for summarized context |
CLAUDE_CODE_AUTO_COMPACT_BUFFER_TOKENS |
Override auto-compact buffer size |
CLAUDE_CODE_ADVISOR_MODEL |
Set the advisor/reviewer model (provider-agnostic) |
CLAUDE_CODE_SUBAGENT_MAX_CONTEXT_TOKENS |
Override max context window size for subagents |
CLAUDE_CODE_SUBAGENT_BUFFER_TOKENS |
Override auto-compact buffer size for subagents |
CLAUDE_CODE_SUBAGENT_SUMMARY_OUTPUT_TOKENS |
Override summary output token reservation for subagents |
CLAUDE_CODE_ADVISOR_MAX_CONTEXT_TOKENS |
Override max context window size for the advisor tool |
CLAUDE_CODE_ADVISOR_BUFFER_TOKENS |
Override auto-compact buffer size for the advisor tool |
CCLITE_TURBO |
Set to 1 (or use --turbo / /turbo) to enable turbo high-concurrency mode |
CCLITE_TURBO_HEDGES |
Total hedged attempts per request including the original (default 2, max 4) |
CCLITE_TURBO_HEDGE_DELAY_MS |
Stagger between hedged attempt starts (default 8000) |
CCLITE_TURBO_ATTEMPT_TIMEOUT_MS |
Abort an attempt that shows no progress within this window; the race continues (0 = off) |
CCLITE_TURBO_INCLUDE_LOCAL |
Set to 1 to hedge against local servers (Ollama/LM Studio) too — usually pointless, it just doubles GPU load |
CCLITE_SEMANTIC_RERANK |
Set to 1 to rerank SearXNG results by local-model semantic similarity (offline, no API cost) |
Some providers are slow per request but tolerate high concurrency. Turbo mode
exploits that with hedged requests: the same request is fired up to
CCLITE_TURBO_HEDGES times, staggered CCLITE_TURBO_HEDGE_DELAY_MS apart. The
first stream to produce real content wins; every other attempt is aborted.
On a slow relay this cuts first-token tail latency dramatically — at the cost
of extra billed input tokens for the aborted duplicates.
Concurrency is dynamic, not a static switch:
- An AIMD governor (TCP-style congestion control) starts at the configured ceiling and adapts per request: sustained successes raise the allowance, 429s/server errors halve it and open a cooldown, and local event-loop saturation (sampled every second) does the same. The floor is a single request, so under pressure the mode degrades gracefully instead of freezing the CLI.
- The same allowance drives parallel tool/subagent fan-out in turbo mode
(ceiling 20;
CLAUDE_CODE_MAX_TOOL_USE_CONCURRENCYstays authoritative). - Local backends (Ollama/LM Studio) gain nothing from duplicate requests, so
they are skipped unless
CCLITE_TURBO_INCLUDE_LOCAL=1. CCLITE_TURBO_ATTEMPT_TIMEOUT_MSreaps silently-hanging attempts so one dead upstream connection can never stall the race.
Check live behavior with /turbo status (shows current allowance and measured
event-loop lag). Applies to OpenAI-compatible providers (Chat Completions and
Responses transports); the native Anthropic path is untouched.
turbo + bypassPermissions compose freely: bypass only changes the
permission layer, turbo only the API/tool layers — e.g.
CCLITE_TURBO=1 cclite-bypass -p "..." runs both at once.
cclite --turbo # session-long turbo mode
# or toggle inside a session: /turbo on | /turbo off | /turbo status| CLAUDE_CODE_ADVISOR_SUMMARY_OUTPUT_TOKENS | Override summary output token reservation for the advisor tool |
Autocompact buffer = CLAUDE_CODE_SUMMARY_OUTPUT_TOKENS + CLAUDE_CODE_AUTO_COMPACT_BUFFER_TOKENS
Subagents (spawned via the Agent tool) and the Advisor tool each run their own
query loops with independent context management. By default they inherit the same
context window, auto-compact buffer, and summary output reservation as the main
agent. You can override these independently using env vars or settings.json.
Env vars (highest priority):
# Subagent overrides
export CLAUDE_CODE_SUBAGENT_MAX_CONTEXT_TOKENS=100000
export CLAUDE_CODE_SUBAGENT_BUFFER_TOKENS=5000
export CLAUDE_CODE_SUBAGENT_SUMMARY_OUTPUT_TOKENS=8000
# Advisor overrides
export CLAUDE_CODE_ADVISOR_MAX_CONTEXT_TOKENS=150000
export CLAUDE_CODE_ADVISOR_BUFFER_TOKENS=10000
export CLAUDE_CODE_ADVISOR_SUMMARY_OUTPUT_TOKENS=16000settings.json (persistent config):
{
"subagentContextWindow": 100000,
"subagentBufferTokens": 5000,
"subagentSummaryOutputTokens": 8000,
"advisorContextWindow": 150000,
"advisorBufferTokens": 10000,
"advisorSummaryOutputTokens": 16000
}Priority chain (each parameter independently):
- context-specific env var (e.g.
CLAUDE_CODE_SUBAGENT_MAX_CONTEXT_TOKENS) - context-specific
settings.jsonfield (e.g.subagentContextWindow) - general env var (e.g.
CLAUDE_CODE_MAX_CONTEXT_TOKENS) - model default
Context window overrides use min() semantics — they only cap downward, never
expand beyond what the model supports.
Calculation reminder:
effectiveWindow = min(modelContextWindow, contextWindowOverride, AUTO_COMPACT_WINDOW)
- summaryOutputTokens
compactThreshold = effectiveWindow - bufferTokens
totalBuffer = summaryOutputTokens + bufferTokens (default: 20000 + 13000 = 33000)
Use OPENAI_MODEL when you want to pin the whole session to one exact model. Use ANTHROPIC_DEFAULT_OPUS_MODEL, ANTHROPIC_DEFAULT_SONNET_MODEL, and ANTHROPIC_DEFAULT_HAIKU_MODEL when you want aliases such as opus, sonnet, and haiku to resolve to your own provider-specific model IDs.
- Codex aliases such as
codexplanandcodexsparkalso use the unified Responses transport, but through the ChatGPT Codex auth/backend path. - Other OpenAI-compatible backends stay on Chat Completions by default for maximum compatibility unless you explicitly select Responses.
- Official OpenAI requests that include reasoning still use the Responses API.
- Set
OPENAI_API_MODE=responsesto force/responseson providers that support it, orOPENAI_API_MODE=chat_completionsto force the legacy path.
OpenAI (Responses API, explicit):
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=sk-...
export OPENAI_MODEL=gpt-5.4
export OPENAI_API_MODE=responsesOpenAI ChatGPT:
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=sk-...
export OPENAI_MODEL=gpt-4o
export OPENAI_API_MODE=chat_completionsOpenAI ChatGPT Codex:
export CLAUDE_CODE_USE_OPENAI=1
export CODEX_API_KEY=eyJ...
export CHATGPT_ACCOUNT_ID=account_...
export OPENAI_MODEL=codexplanOpenRouter:
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=sk-or-v1-...
export OPENAI_BASE_URL=https://openrouter.ai/api/v1
export OPENAI_MODEL=openai/gpt-5.4DeepSeek:
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=sk-...
export OPENAI_BASE_URL=https://api.deepseek.com/v1
export OPENAI_MODEL=deepseek-chatLLaMA.CPP (local):
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_BASE_URL=http://localhost:1234/v1
export OPENAI_MODEL=your-model-name
export CLAUDE_CODE_MAX_CONTEXT_TOKENS=65536
export CLAUDE_CODE_SUMMARY_OUTPUT_TOKENS=12000
export CLAUDE_CODE_AUTO_COMPACT_BUFFER_TOKENS=4000Ollama (local):
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_BASE_URL=http://localhost:1234/v1
export OPENAI_MODEL=llama3.3LM Studio (local):
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_BASE_URL=http://localhost:1234/v1
export OPENAI_MODEL=your-model-nameAzure OpenAI:
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=your-azure-key
export OPENAI_BASE_URL=https://your-resource.openai.azure.com/openai/deployments/your-deployment
export OPENAI_MODEL=gpt-4o
export AZURE_OPENAI_API_VERSION=2024-12-01-previewYou can route CC-lite's built-in WebSearch tool through your own SearXNG instance by setting one environment variable:
export CLAUDE_CODE_SEARXNG_BASE_URL=http://localhost:8888/When this variable is set, WebSearch no longer relies on the provider's server-side web search. Instead, CC-lite calls your SearXNG instance directly at:
GET {CLAUDE_CODE_SEARXNG_BASE_URL}/search?q=<query>&format=json
Notes:
- This only changes the
WebSearchtool.WebFetchstill fetches page content directly. allowed_domainsandblocked_domainsare still supported, but filtering is applied locally after SearXNG returns results.- If
CLAUDE_CODE_SEARXNG_BASE_URLis unset, CC-lite falls back to the default provider behavior. - Optional: set
CCLITE_SEMANTIC_RERANK=1and hits are reranked by semantic similarity to the query using the local advisor embedding model — offline, free, best-effort (any failure keeps the original order).
The original Claude Code source is the property of Anthropic. This fork exists because the source was publicly exposed through their npm distribution. Use at your own discretion.