What you would like to be added?
Add a Grafana dashboard for TrainJob GPU utilization using dcgm-exporter metrics. This would complement the controller health dashboard added in #3653 by providing visibility into GPU resource usage during training.
The dashboard could be modeled after the NVIDIA dcgm-exporter dashboard and adapted for TrainJob workloads.
It should follow the same grafanaDashboard selection pattern introduced in #3653, allowing users to enable it individually or via grafanaDashboard.defaultEnabled.
Why is this needed?
GPU utilization is a key metric for ML training workloads. Users running TrainJobs on GPU nodes need visibility into GPU memory usage, compute utilization, and temperature to optimize training performance and cost. Currently, the Grafana dashboard in #3653 only covers controller health and reconciliation metrics.
Ref: #3653 (comment)
What you would like to be added?
Add a Grafana dashboard for TrainJob GPU utilization using dcgm-exporter metrics. This would complement the controller health dashboard added in #3653 by providing visibility into GPU resource usage during training.
The dashboard could be modeled after the NVIDIA dcgm-exporter dashboard and adapted for TrainJob workloads.
It should follow the same
grafanaDashboardselection pattern introduced in #3653, allowing users to enable it individually or viagrafanaDashboard.defaultEnabled.Why is this needed?
GPU utilization is a key metric for ML training workloads. Users running TrainJobs on GPU nodes need visibility into GPU memory usage, compute utilization, and temperature to optimize training performance and cost. Currently, the Grafana dashboard in #3653 only covers controller health and reconciliation metrics.
Ref: #3653 (comment)