NemoRL multinode training example - #89
Conversation
| volumes=modal_volumes, | ||
| secrets=[ | ||
| modal.Secret.from_name("huggingface-secret"), | ||
| modal.Secret.from_name("wandb-secret"), |
There was a problem hiding this comment.
🟡 STYLE_GUIDE violation: wandb-secret is required but should be optional
The STYLE_GUIDE.md states: "wandb-secret should be used for Weights & Biases. This secret should always be optional i.e. the script should work without it." The train function at nemo-rl/modal_train.py:151-152 includes modal.Secret.from_name("wandb-secret") as a required secret (the default for from_name is required=True). This means running modal run modal_train.py::train will fail at deploy time if the user hasn't configured a wandb-secret Modal secret, even though wandb logging is conceptually optional in NeMo-RL (it's just an override in each config).
| modal.Secret.from_name("wandb-secret"), | |
| modal.Secret.from_name("wandb-secret", required=False), |
Was this helpful? React with 👍 or 👎 to provide feedback.
| try: | ||
| table = placement_group_table(placement_group) | ||
| return table.get("bundles_to_node_id", {}).get(bundle_index) | ||
| except Exception: | ||
| return None |
There was a problem hiding this comment.
🚩 _bundle_node_id dict key type may not match bundle_index int
In _bundle_node_id at nemo-rl/modal_helpers/run_grpo_multinode.py:53, table.get("bundles_to_node_id", {}).get(bundle_index) uses bundle_index as an int. Depending on the Ray version, placement_group_table may return bundles_to_node_id with string keys (from protobuf/JSON serialization). If the keys are strings like "0", "1", the int lookup would always return None, causing the NCCL_HOSTID patch to silently never apply. The function is wrapped in a broad try/except returning None, so it would fail silently. This is hard to verify without the exact Ray version in the container image, but if this lookup fails, multinode training could hit 'Duplicate GPU detected' NCCL errors.
Was this helpful? React with 👍 or 👎 to provide feedback.
This adds a self-contained launcher for running multinode training using NemoRL on Modal, covering both single-node and multi-node (RDMA/EFA) clusters.
The organization is taken from the Slime example, with a small config system (
configs/base.py) where each experiment is a Python module exposing a modal object and nemo_rl object (which run script to use, which base YAML from NemoRL, and a dict of overrides). Similar to the Slime example,modal_train.pyturns that into a Modal app withdownload_model,download_data, andtrainentry points, starting a Ray cluster across the allocated nodes and running the NemoRL driver on the head node w/ the HF cache and checkpoints mounted as volumes.Tested examples: GRPO on OpenMathInstruct-2 math task: Qwen2.5-1.5B (single and 2-node DP), Llama-3.1-8B two-node, Qwen3-8B, and a Nemotron-Nano-v3-30B-A3B MoE FSDP + EP two-node config. README contains setup + launch instructions