Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
29 commits
Select commit Hold shift + click to select a range
1820642
slime: plumb ModalConfig memory/cloud/region into Modal Function
nanjiangwill Apr 17, 2026
54852c2
remove miles v1
nanjiangwill Apr 10, 2026
dda1ce8
Add miles v2 implementation plan and draft
nanjiangwill Apr 10, 2026
fdf6553
Add miles v2 Modal launcher with Qwen3-4B LoRA smoke config
nanjiangwill Apr 10, 2026
499a8e2
Fix dataset sources and improve README for miles launcher
nanjiangwill Apr 10, 2026
b5a9606
Make wandb-secret conditional on use_wandb config flag
nanjiangwill Apr 10, 2026
35f67aa
Enable WandB in miles smoke-test config
nanjiangwill Apr 10, 2026
1e13ee4
Revert conditional wandb-secret; always attach like slime
nanjiangwill Apr 10, 2026
8ae7daa
add K2.5 Modal configs (smoke + LoRA) and bridge test script
nanjiangwill Apr 16, 2026
7d8fd5d
lean kimi_k25_fullparam_smoke: remove pynccl patch, stale INT4 QAT stub
nanjiangwill Apr 16, 2026
dd6c8be
kimi_k25_lora: phased target_modules (MLA only), enable INT4 QAT
nanjiangwill Apr 16, 2026
b1c8e6b
kimi_k25_lora: cover attention + MLP (not attention-only)
nanjiangwill Apr 16, 2026
5d5144c
kimi 2.5 working full param
nanjiangwill Apr 20, 2026
f1a2ceb
working kimi 2.5 lora
nanjiangwill Apr 20, 2026
138aa31
working kimi 2.5 lora
nanjiangwill Apr 20, 2026
2dbe593
cleanup
nanjiangwill Apr 20, 2026
1180b38
cleanup
nanjiangwill Apr 20, 2026
f4e3465
cleanup
nanjiangwill Apr 20, 2026
f4904ae
install megatron bridge
nanjiangwill Apr 20, 2026
e74993a
cleanup
nanjiangwill Apr 20, 2026
5c32ddb
cleanup
nanjiangwill Apr 20, 2026
53ffe81
working kimi 2.5 lora mla without oom
nanjiangwill Apr 21, 2026
fa4c243
:bug: Fix imports to be global not relative for patches
joyliu-q Apr 20, 2026
d5c9bd3
:tada: Add path resolution
joyliu-q Apr 20, 2026
b775367
:fire: Remove some agent files
joyliu-q Apr 20, 2026
d4768ce
:art: Directory detection
joyliu-q Apr 20, 2026
5b29802
:new: v0
joyliu-q Apr 21, 2026
a4e3a61
:sparkles: Add ModelConfiguration base class and Kimi smoke test
joyliu-q Apr 21, 2026
2914c58
chore: gitignore .claude/ local state
joyliu-q Apr 21, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -176,6 +176,26 @@ cython_debug/

# experimental
slime/test_configs
slime/configs/glm47_355b_a32b_noncolocate.py
slime/configs/glm47_355b_a32b_noncolocate_delta_compression.py
slime/configs/glm47_355b_a32b_noncolocate_delta_compression_bitmask.py
slime/configs/glm47_355b_a32b_noncolocate_delta_compression_indices.py
slime/configs/glm47_flash_dapo_noncolocate_dela_compression.py
slime/configs/qwen3_235b_a22b_dapo_non_colocate.py
slime/configs/qwen3_dapo_noncolocate_delta_compression.py
slime/configs/qwen3_dapo_noncolocate_delta_compression_bitmask.py
slime/configs/qwen3_dapo_noncolocate_delta_compression_indices.py
slime/configs/qwen3vl_geo3k_vlm_delta_compression.py

# Ruff cache
.ruff_cache/

# Codex cache
.codex
.codex/

# Humanize plugin state
.humanize/

# Claude Code local state
.claude/
250 changes: 179 additions & 71 deletions miles/README.md
Original file line number Diff line number Diff line change
@@ -1,121 +1,229 @@
# Miles on Modal
# miles — Modal launcher for Miles training

Run Miles RL training on Modal with recipe files stored under [`recipes/`](./recipes/).
The wrapper handles Modal and Ray orchestration; model and training flags stay in
recipe arg files.
Thin Modal launcher that runs [Miles](https://github.com/radixark/miles) RL training on GPU clusters.

## Prerequisites

- A Modal account with multi-node access.
- A `huggingface-secret` Modal secret containing `HF_TOKEN`.
- Optional: `WANDB_API_KEY` in your local shell for Weights & Biases logging.
- Optional: `modal deploy miles/modal_train.py`. The local entrypoint will try
the deployed `MilesCluster` first and fall back to an ephemeral app if it is
not deployed.
- Modal CLI installed and authenticated
- Set your Modal environment: `export MODAL_ENVIRONMENT=<your-env>`
- Modal secrets:
- `huggingface-secret` — required for `prepare_model` and `prepare_data`
- `wandb-secret` — required only for experiments with `use_wandb = True`

## Prepare Shared Assets
## Running an experiment

Prepare the default GSM8K dataset:
All commands take the experiment name via `EXPERIMENT_CONFIG`. Run from the repo root.

### 1. List available experiments

```bash
modal run miles/modal_train.py::prepare_dataset
modal run miles/modal_train.py::list_configs
```

Download a model for a built-in recipe:
### 2. Prepare model (one-time)

Downloads the experiment's HF checkpoint to the `huggingface-cache` volume and applies any experiment-specific model fixes.

```bash
modal run miles/modal_train.py::download_model --recipe qwen3-30b-a3b-lora
EXPERIMENT_CONFIG=qwen3_4b_lora_smoke modal run miles/modal_train.py::prepare_model
```

Or download any model directly:
### 3. Prepare data (one-time)

Downloads and preprocesses the training dataset to the `miles-data` volume.
Only required if the experiment defines a `prepare_data()` function (see [Adding an experiment](#adding-an-experiment)).

```bash
modal run miles/modal_train.py::download_model --model-id Qwen/Qwen3-30B-A3B
EXPERIMENT_CONFIG=qwen3_4b_lora_smoke modal run miles/modal_train.py::prepare_data
```

## Recipes
### 4. Convert checkpoint (one-time, raw mode only)

List the available recipes:
Converts the HF checkpoint to `torch_dist` format. Only required when `megatron_to_hf_mode = "raw"`.
Skip this step if using bridge mode.

```bash
modal run miles/modal_train.py --list-recipes
EXPERIMENT_CONFIG=qwen3_4b_lora_smoke modal run modal_train.py::convert_checkpoint
```

Recommended starting points:
### 5. Run training

- `qwen3-30b-a3b-lora`: default Qwen3 recipe.
- `qwen3-30b-a3b-lora-fewstep`: smallest end-to-end Qwen3 validation recipe.
- `qwen3-30b-a3b-experts-lora`: explicit expert-target variant.
- `qwen3-30b-a3b-experts-fewstep`: trimmed expert-target validation recipe.
- `qwen25-0p5b-lora`: small smoke test.
```bash
EXPERIMENT_CONFIG=qwen3_4b_lora_smoke modal run -d modal_train.py::train
```

Testing and debug recipes live under [`recipes/tests/`](./recipes/tests).
Use `-d` (detached) to keep training running after you close your terminal.

## Train
## Docker Image

Set `MILES_N_NODES` in the same shell invocation as `modal run`.
Uses `radixark/miles:dev-202604101227` as the base image.

Single-node Qwen3 few-step validation:
## Architecture

```bash
MILES_N_NODES=1 modal run miles/modal_train.py --recipe qwen3-30b-a3b-lora-fewstep
```
Each experiment defines two module-level objects in `configs/<name>.py`:

Single-node Qwen3 default recipe:
- `modal` — `ModalConfig` instance (GPU type, dev overlays, image patches)
- `miles` — `MilesConfig` subclass instance (all Miles CLI arguments)

```bash
MILES_N_NODES=1 modal run miles/modal_train.py --recipe qwen3-30b-a3b-lora
`MilesConfig` attributes are automatically converted to CLI flags:
- `lora_rank = 64` → `--lora-rank 64`
- `colocate = True` → `--colocate`
- `colocate = False` → omitted

### Special Fields

These `MilesConfig` fields are **not** passed as CLI args:

| Field | Purpose |
|-------|---------|
| `environment` | Injected into the Ray job runtime env |
| `async_mode` | Selects `train_async.py` vs `train.py` |
| `miles_model_script` | Shell script sourced for `MODEL_ARGS` (e.g., `scripts/models/qwen3-4B.sh`) |

### Volumes

| Volume | Mount Path | Purpose |
|--------|-----------|---------|
| `huggingface-cache` | `/root/.cache/huggingface` | Model checkpoints |
| `miles-data` | `/data` | Training datasets |
| `miles-checkpoints` | `/checkpoints` | Converted checkpoints |

## Adding an experiment

### 1. Create the config file

Create `configs/<your_experiment>.py`. Each config file must expose two module-level instances:
- `modal` — a `ModalConfig` instance (GPU type, image patches)
- `miles` — a `MilesConfig` subclass instance (all Miles training arguments)

```python
from configs.base import ModalConfig, MilesConfig, DATA_PATH

modal = ModalConfig(gpu="H200")


class _Miles(MilesConfig):
# Launcher instructions (not passed to Miles CLI)
miles_model_script = "scripts/models/qwen3-4B.sh" # sources MODEL_ARGS
async_mode = False

# Model
hf_checkpoint = "Qwen/Qwen3-4B"
megatron_to_hf_mode = "bridge" # or "raw" (requires convert_checkpoint)

# Infrastructure
actor_num_nodes = 1
actor_num_gpus_per_node = 4
colocate = True

# Data
prompt_data = f"{DATA_PATH}/my_dataset/train.jsonl"
input_key = "prompt"
label_key = "label"
rm_type = "deepscaler"

# ... all other Miles args as snake_case attributes


miles = _Miles()
```

Single-node expert-target follow-up:
Every attribute on `_Miles` (except `environment`, `async_mode`, `miles_model_script`) is forwarded to
Miles as a CLI argument: `field_name` → `--field-name`. See `configs/base.py` for full rules.

```bash
MILES_N_NODES=1 modal run miles/modal_train.py --recipe qwen3-30b-a3b-experts-fewstep
### 2. Add `prepare_model()` / `prepare_data()` methods (if needed)

Override `prepare_model()` if your experiment needs model-specific preparation beyond a plain HF download.
The base implementation already calls `snapshot_download(self.hf_checkpoint)`.

```python
class _Miles(MilesConfig):
...
def prepare_model(self) -> None:
super().prepare_model()
# apply model-specific local patches if needed
```

Non-colocated Qwen3 validation:
If your experiment needs to download or preprocess a dataset, override `prepare_data()` on `_Miles`.
It runs inside the Modal container with the `miles-data` volume mounted at `DATA_PATH`.

```python
class _Miles(MilesConfig):
...
def prepare_data(self) -> None:
import os
from huggingface_hub import snapshot_download

os.makedirs(f"{DATA_PATH}/my_dataset", exist_ok=True)
snapshot_download(
repo_id="org/my-dataset",
repo_type="dataset",
local_dir=f"{DATA_PATH}/my_dataset",
)
```

If `prepare_data()` is not overridden, `prepare_data` will raise `NotImplementedError` — simply skip that step.

### 3. Run the workflow

`EXPERIMENT_CONFIG` is the config filename without `.py`:

```bash
MILES_N_NODES=2 modal run miles/modal_train.py \
--recipe qwen3-30b-a3b-lora-fewstep \
--no-colocate \
--actor-nodes 1 \
--allow-cluster-mismatch
EXPERIMENT_CONFIG=my_experiment modal run miles/modal_train.py::prepare_model
EXPERIMENT_CONFIG=my_experiment modal run miles/modal_train.py::prepare_data # if prepare_data() defined
EXPERIMENT_CONFIG=my_experiment modal run miles/modal_train.py::convert_checkpoint # if megatron_to_hf_mode = "raw"
EXPERIMENT_CONFIG=my_experiment modal run -d miles/modal_train.py::train
```

Small smoke test:
No registration step needed — the launcher discovers configs automatically from the `configs/` directory.

```bash
MILES_N_NODES=1 modal run miles/modal_train.py --recipe qwen25-0p5b-lora
## YAML config fields

`eval_config`, `custom_config_path`, and `sglang_config` normally take file paths in Miles.
In Python configs you can write them as inline dicts — the launcher materializes them to temp YAML files automatically:

```python
class _Miles(MilesConfig):
eval_config = {
"eval": {
"defaults": {"max_response_len": 16384},
"datasets": [
{"name": "aime", "path": "/data/aime.jsonl", "rm_type": "deepscaler"},
],
}
}
```

## Ad Hoc Runs
## JSON config fields

You can launch without a predefined recipe by passing args directly:
`train_env_vars`, `apply_chat_template_kwargs`, and `multimodal_keys` are parsed by Miles with `json.loads()`.
If set as dicts in Python configs, the launcher serializes them with `json.dumps()` automatically.

```bash
MILES_N_NODES=1 modal run miles/modal_train.py \
--model-id Qwen/Qwen3-30B-A3B \
--args-file miles/recipes/qwen3-30b-a3b-lora.args \
--extra-args "--train-samples 8 --eval-interval 1" \
--run-name qwen3-adhoc
## Dev overlay

To test local Miles changes without rebuilding the image, set `local_miles` in your `ModalConfig`:

```python
modal = ModalConfig(
gpu="H200",
local_miles="/path/to/your/miles",
)
```

## Useful Options
## Applying patches to the image

To inject local patch files into the image (e.g. to patch SGLang), use `patch_files` and `image_run_commands`:

- `--dry-run`: print the assembled Miles command without launching a job.
- `--args` / `--args-file`: provide the base Miles CLI args.
- `--extra-args` / `--extra-args-file`: append overrides to a recipe or ad hoc run.
- `--custom-config`: pass a YAML override file through to Miles.
- `--run-name`: override the checkpoint subdirectory name.
- `--allow-cluster-mismatch`: bypass recipe node-count checks.
- `USE_LOCAL_MILES=/path/to/miles`: overlay a local Miles checkout.
- `MILES_IMAGE=radixark/miles:...`: override the pinned container image.
```python
from pathlib import Path

## Notes
_HERE = Path(__file__).resolve().parent

modal = ModalConfig(
gpu="H200",
patch_files=[str(_HERE.parent / "patches" / "sglang_fix.patch")],
image_run_commands=["cd /sgl-workspace/sglang && git apply /tmp/sglang_fix.patch"],
)
```

- The default Qwen3 recipes use standard all-layer LoRA over
`linear_qkv`, `linear_proj`, `linear_fc1`, and `linear_fc2`.
- Start with the few-step recipes when validating a new environment.
- Modal-specific runtime compatibility patches live in
[`modal_patches/sitecustomize.py`](./modal_patches/sitecustomize.py).
Each file in `patch_files` is added to the image at `/tmp/<filename>`.
15 changes: 15 additions & 0 deletions miles/configs/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
import importlib
from pathlib import Path

_CONFIGS_DIR = Path(__file__).parent
_SKIP = {"base", "__init__", "model_configuration"}


def get_module(name: str):
try:
return importlib.import_module(f"configs.{name}")
except ModuleNotFoundError:
available = sorted(
f.stem for f in _CONFIGS_DIR.glob("*.py") if f.stem not in _SKIP
)
raise ValueError(f"Unknown config {name!r}. Available: {available}")
Loading