GPU / MIG Operations Dashboard
Detailed Grafana-style visual monitoring combined with PSI PULSE L1 workflow and fleet context.
Aggregate
1,024 servers • 8,192 GPUs • 3 hallsPer GPU
DCGM real-time metrics8,192 GPUs
Per MIG Profile
Memory, utilization and allocation| Hostname | Hall / Rack | GPU ID | MIG Profile | GPU Instance | Compute Instance | Namespace | Pod | GPU Util | Memory | Health |
|---|---|---|---|---|---|---|---|---|---|---|
| h4-r12-n031 | H4 / R12 | 0 | 1g.24gb | 13 | 0 | ai-inference | model-a-7f9 | 81% | 20.1 GB | Healthy |
| h4-r12-n031 | H4 / R12 | 1 | 2g.48gb | 3 | 0 | ai-training | train-llm-02 | 94% | 43.6 GB | Healthy |
| h3-r07-n018 | H3 / R07 | 0 | 2g.48gb | 4 | 0 | vision-prod | vision-24 | 37% | 31.2 GB | Thermal watch |
| h3-r08-n004 | H3 / R08 | 3 | 1g.24gb | 9 | 0 | ai-inference | embed-03 | 0% | 3.2 GB | XID detected |
Error & Reliability
PCIe, XID and ECCL1 Operations Queue
Correlated by PSI PULSECollect DCGM diagnostics and BMC events. Do not reset GPU until workload drain approval.
Validate optics and BER counters on Leaf-07 and prepare OEM evidence bundle.
17 idle MIG instances are candidates for scheduler review, not automatic reassignment.