You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: mkdocs/docs/examples/training/miles.md
+44-33Lines changed: 44 additions & 33 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,22 +1,32 @@
1
1
---
2
2
title: Miles
3
-
description: RL-fine-tune Qwen2.5-32B with Miles, SGLang, Megatron-LM, and Ray across two 8xH100 nodes
3
+
description: RLfine-tuning Qwen2.5-32B with Miles, SGLang, Megatron-LM, and Ray across two 8xH100 nodes
4
4
---
5
5
6
6
# Miles
7
7
8
8
This example shows how to use `dstack` and [Miles](https://github.com/radixark/miles)
9
-
to RL-fine-tune a 32B language model with [GRPO](https://arxiv.org/abs/2402.03300) across two 8xH100 nodes. Under the hood Miles uses [SGLang](https://github.com/sgl-project/sglang) for high-throughput rollout, [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) for training, and [Ray](https://docs.ray.io/en/latest/) to coordinate the trainer
10
-
and rollout actors across nodes.
9
+
to fine-tune a 32B language model with [GRPO](https://arxiv.org/abs/2402.03300)
10
+
across a multi-node cluster.
11
+
Miles uses [SGLang](https://github.com/sgl-project/sglang) for high-throughput
12
+
rollouts, [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) for training,
13
+
and [Ray](https://docs.ray.io/en/latest/) to coordinate the trainer and rollout
14
+
actors across nodes.
11
15
12
-
Here we fine-tune `Qwen/Qwen2.5-32B-Instruct` on [GSM8K](https://huggingface.co/datasets/openai/gsm8k) dataset.
16
+
Here we fine-tune `Qwen/Qwen2.5-32B-Instruct` on the
Before running a distributed task, make sure to create a fleet with `placement` set to `cluster` (can be a [managed fleet](../../concepts/fleets.md#cluster-placement) or an [SSH fleet](../../concepts/fleets.md#ssh-placement)).
20
+
Before running a distributed task, make sure to create a [fleet](../../concepts/fleets.md)
21
+
with `placement` set to [`cluster`](../../concepts/fleets.md#cluster-placement).
16
22
17
23
## Run a Ray cluster
18
24
19
-
The task below starts a Ray cluster across both nodes and performs the one-time setup on each node. The setup includes: downloading the model, downloading the dataset, and converting the checkpoint to Megatron's `torch_dist` format.
25
+
### Define a configuration
26
+
27
+
The [task](../../concepts/tasks.md) below starts Ray on two nodes and prepares
28
+
each node by downloading the model and dataset, then converting the checkpoint
As long as the `dstack apply` is attached, you can use `localhost:8265` to submit Ray jobs for execution. If `dstack apply` is detached, you can use `dstack attach` to re-attach.
103
+
While `dstack apply` is attached, you can submit Ray jobs through
104
+
`localhost:8265`. If you detach or run from another machine, use
105
+
[`dstack attach`](../../reference/cli/dstack/attach.md) to re-attach and make
106
+
the dashboard port accessible on `localhost`.
91
107
92
108
## Submit Ray jobs
93
109
94
-
Before you can submit Ray jobs, ensure `ray` is installed locally:
110
+
Install `ray` locally before submitting jobs:
95
111
96
112
<div class="termy">
97
113
@@ -102,7 +118,7 @@ $ pip install ray
102
118
</div>
103
119
104
120
The submit script below runs the Miles training job on the Ray cluster. The
105
-
model is sharded across all 8 GPUs per node via tensor parallelism, and SGLang
121
+
model is sharded across all 8 GPUs per node with tensor parallelism, and SGLang
0 commit comments