| slug | Installation |
|---|---|
| title | Server deployment |
| description | Deploy and run the Ultron server (self-hosted) |
This guide explains how to run your own Ultron service. If you only need an assistant to connect to an existing Ultron instance (self-hosted or public), see Agent setup.
git clone https://github.com/modelscope/ultron.git
cd ultron
pip install -e .Core dependencies (see requirements.txt):
| Package | Purpose |
|---|---|
fastapi |
HTTP framework |
uvicorn |
ASGI server |
pydantic |
Validation |
tiktoken |
Token counting |
dashscope |
Text embeddings |
openai |
OpenAI-compatible LLM access |
| Requirement | Notes |
|---|---|
| Python | >= 3.8 |
| OS | Linux / macOS / Windows |
Ultron primarily calls LLM APIs, so a CPU-only machine is enough.
Smart ingestion and LLM-based classification need OpenAI-compatible LLM config. Recommended: set ULTRON_LLM_PROVIDER, ULTRON_MODEL, ULTRON_BASE_URL, and ULTRON_API_KEY in ~/.ultron/.env. Embeddings default to DASHSCOPE_API_KEY; to use another model, set ULTRON_EMBEDDING_BACKEND to local (local sentence-transformers) or openai (any OpenAI-compatible /embeddings endpoint such as Zhipu embedding-3, OpenAI, vLLM/TEI, configured via ULTRON_EMBEDDING_BASE_URL/ULTRON_EMBEDDING_API_KEY).
You can also export in the shell for one-off debugging:
export ULTRON_LLM_PROVIDER="openai"
export ULTRON_MODEL="gpt-5"
export ULTRON_BASE_URL="https://api.openai.com/v1"
export ULTRON_API_KEY="your-api-key"Using ~/.ultron/.env requires python-dotenv; see the repo root .env.example. Importing ultron only auto-loads ~/.ultron/.env (see Configuration). In systemd, Docker, etc., inject the same variable names.
Other optional variables (full ULTRON_* list in Configuration); constructor arguments to UltronConfig(...) override the environment.
| Variable | Description | Default |
|---|---|---|
ULTRON_EMBEDDING_MODEL |
TextEmbedding model name | text-embedding-v4 |
ULTRON_EMBEDDING_BACKEND |
Embedding backend (dashscope, openai, or local) |
dashscope |
ULTRON_EMBEDDING_DIMENSION |
Vector dimension | 1024 |
ULTRON_EMBEDDING_BASE_URL |
openai backend OpenAI-compatible /embeddings root URL (falls back to ULTRON_BASE_URL) |
"" |
ULTRON_EMBEDDING_API_KEY |
openai backend API key (falls back to ULTRON_API_KEY / OPENAI_API_KEY) |
"" |
ULTRON_LLM_PROVIDER |
OpenAI-compatible provider selector | dashscope |
ULTRON_MODEL |
LLM for smart ingestion and extraction | qwen3.6-flash |
ULTRON_MEMORY_CATEGORY_MODEL |
LLM for memory type classification | qwen3.6-flash |
ULTRON_SKILL_CATEGORY_MODEL |
LLM for skill taxonomy | qwen3.6-flash |
ULTRON_BASE_URL |
OpenAI-compatible API base | https://dashscope.aliyuncs.com/compatible-mode/v1 |
ULTRON_API_KEY |
LLM API key | "" |
| Variable | Description | Default |
|---|---|---|
ULTRON_LLM_MAX_INPUT_TOKENS |
Max user text tokens | 200000 |
ULTRON_LLM_PROMPT_RESERVE_TOKENS |
Reserved tokens for system prompt | 8192 |
ULTRON_LLM_TOKEN_COUNT_ENCODING |
tiktoken encoding name | cl100k_base |
ULTRON_LLM_REQUEST_TIMEOUT |
HTTP read timeout (seconds) | 600 |
ULTRON_LLM_MAX_RETRIES |
Retries after failure | 2 |
ULTRON_LLM_RETRY_BASE_DELAY |
Base backoff (seconds) | 1.0 |
| Variable | Description | Default |
|---|---|---|
ULTRON_MEMORY_MERGE_MAX_FIELD_TOKENS |
Max tokens per field when merging | 8192 |
| Variable | Description | Default |
|---|---|---|
ULTRON_L0_MAX_TOKENS |
Max tokens for L0 summary in results | 64 |
ULTRON_L1_MAX_TOKENS |
Max tokens for L1 snippet in results | 256 |
| Variable | Description | Default |
|---|---|---|
ULTRON_DATA_DIR |
Data root (~ expanded) |
~/.ultron |
ULTRON_DB_NAME |
SQLite filename | ultron.db |
ULTRON_HOT_MAX_ENTRIES |
HOT tier max entries | 500 |
ULTRON_WARM_MAX_ENTRIES |
WARM tier max entries | 1000 |
ULTRON_HOT_PERCENTILE |
HOT percentile | 10 |
ULTRON_DEDUP_SIMILARITY_THRESHOLD |
Cosine threshold for near-duplicate | 0.85 |
ULTRON_ENABLE_INTENT_ANALYSIS |
Intent analysis before search (0 off) |
1 |
ULTRON_MEMORY_SEARCH_LIMIT |
Default memory search limit | 10 |
ULTRON_SKILL_SEARCH_LIMIT |
Default skill search limit | 5 |
ULTRON_DECAY_INTERVAL_HOURS |
Decay job interval (hours) | 6.0 |
ULTRON_DECAY_ALPHA |
Freshness coefficient | 0.05 |
ULTRON_COLD_TTL_DAYS |
COLD retention days (0 = never delete) |
30 |
| Variable | Description | Default |
|---|---|---|
ULTRON_LOG_LEVEL |
Log level | INFO |
ULTRON_RESET_TOKEN |
Auth token for /reset (unset disables) |
none |
More fields (async embedding queue, hot summary caps, etc.) are documented in Configuration.
Important: one
ULTRON_DATA_DIRcan only use one embedding backend/model combination. Before switching embedding backend or model, runreset_all(); otherwise startup validation will fail to prevent mixed vectors and retrieval anomalies.
uvicorn ultron.server:app --host 0.0.0.0 --port 9999
# Default http://0.0.0.0:9999from ultron import Ultron
ultron = Ultron()
# ...Default UltronConfig.data_dir is ~/.ultron/. The library uses:
| Path | Purpose |
|---|---|
ultron.db |
SQLite database |
skills/ |
Skill content |
archive/ |
Archived skills |
models/ |
Local model cache |
When you start ultron.server:app with uvicorn, structured JSON logs go to ~/.ultron/logs/ (ultron.log and rotated backups).
Override paths with ULTRON_DATA_DIR, Ultron(config=UltronConfig(data_dir=...)), or Ultron(data_dir=...).
After the server is up, follow Agent setup to connect your assistant to Ultron.