@@ -219,27 +219,27 @@ description: Quick guide to creating fleets and submitting runs
219219
220220 ```yaml
221221 type: service
222- name: llama31-service
223-
224- # If `image` is not specified, dstack uses its default image
225- python: "3.11"
226- #image: dstackai/base:py3.13-0.7-cuda-12.1
227-
228- # Required environment variables
229- env:
230- - HF_TOKEN
222+ name: qwen36-service
223+
224+ image: vllm/vllm-openai:v0.19.1
225+
231226 commands:
232- - pip install vllm
233- - vllm serve meta-llama/Meta-Llama-3.1-8B-Instruct --max-model-len 4096
234- # Expose the vllm server port
227+ - |
228+ vllm serve Qwen/Qwen3.6-27B \
229+ --host 0.0.0.0 \
230+ --port 8000 \
231+ --max-model-len 32768 \
232+ --reasoning-parser qwen3
233+ # Expose the vLLM server port
235234 port: 8000
236235
237236 # Specify a name if it's an OpenAI-compatible model
238- model: meta-llama/Meta-Llama-3.1-8B-Instruct
239-
237+ model: Qwen/Qwen3.6-27B
238+
240239 # Required resources
241240 resources:
242- gpu: 24GB
241+ shm_size: 16GB
242+ gpu: H100
243243 ```
244244
245245 </div>
@@ -249,22 +249,20 @@ description: Quick guide to creating fleets and submitting runs
249249 <div class="termy">
250250
251251 ```shell
252- $ HF_TOKEN=...
253252 $ dstack apply -f service.dstack.yml
254-
255- # BACKEND REGION INSTANCE RESOURCES SPOT PRICE
256- 1 aws us-west-2 g5.4xlarge 16xCPU, 64GB, 1xA10G (24GB) yes $0.22
257- 2 aws us-east-2 g6.xlarge 4xCPU, 16GB, 1xL4 (24GB) yes $0.27
258- 3 gcp us-west1 g2-standard-4 4xCPU, 16GB, 1xL4 (24GB) yes $0.27
259-
260- Submit the run llama31-service? [y/n]: y
261-
262- Provisioning `llama31-service`...
253+
254+ # BACKEND REGION INSTANCE RESOURCES SPOT PRICE
255+ 1 nebius eu-north1 gpu-h100-sxm 16xCPU, 250GB, 1xH100 (80GB) no $2.95
256+ 2 runpod US-CA-2 NVIDIA H100 80GB HBM3 64xCPU, 1004GB, 1xH100 (80GB) no $2.99
257+
258+ Submit the run qwen36-service? [y/n]: y
259+
260+ Provisioning `qwen36-service`...
263261 ---> 100%
264262
265263 Service is published at:
266- http://localhost:3000/proxy/services/main/llama31 -service/
267- Model meta-llama/Meta-Llama-3.1-8B-Instruct is published at:
264+ http://localhost:3000/proxy/services/main/qwen36 -service/
265+ Model Qwen/Qwen3.6-27B is published at:
268266 http://localhost:3000/proxy/models/main/
269267 ```
270268
0 commit comments