GPU Cluster Monitoring (GCM): Large-Scale AI Research Cluster Monitoring
-
Updated
Aug 31, 2026 - Python
GPU Cluster Monitoring (GCM): Large-Scale AI Research Cluster Monitoring
GPU Observability with workload attribution. One OTLP agent per node ties hardware metrics (NVIDIA, AMD, Intel Gaudi) to the K8s pod or Slurm job burning the GPU.
Hands-on GPU/HPC infrastructure operations: K8s GPU scheduling, HAMi sharing, Slurm, observability & vLLM inference. Learn it free on a laptop; validate on one cheap GPU.
top-like TUI for per-pod GPU usage in Kubernetes (zero cluster footprint)
Simulate NVIDIA GPUs for testing. 7 behavior profiles, scale to 1000+ GPUs, Docker-ready Prometheus exporter using DCGM
Single-file interactive study tool for the NVIDIA NCP-AI Operations exam: quiz, flashcards, guided labs, and a stateful BCM/Slurm/Kubernetes/DCGM command sandbox.
End-to-end observability for disaggregated LLM inference on EKS — DCGM metrics, KEDA autoscaling on GPU signals, per-namespace cost attribution, multi-agent OTel tracing, and Istio mTLS. Reference implementation for the OpenTelemetry AI Inference Platform blueprint.
kubectl plugin that compares requested GPU resources against DCGM Exporter utilization metrics and generates rightsizing recommendations with projected monthly cost savings. Supports nvidia.com/gpu and amd.com/gpu — the gap VPA leaves open.
GPU-native agent-swarm orchestration for the NVIDIA AI stack — NeMo, NIM, Triton, DCGM, NGC, NIXL, OpenShell. Spawn GPU-pinned agent teams across DGX/HGX nodes with NVLink-aware scheduling, task DAGs, adaptive scheduling, and full observability.
Detects and reclaims wasted GPU allocations on Slurm clusters — automated capacity recovery with guardrails that fail safe when telemetry is stale.
Automated acceptance toolkit for Linux deep learning GPU servers
One-command NVIDIA DGX Spark (GB10) cluster GPU monitoring with Grafana + Prometheus + DCGM + node_exporter + vLLM. Track GPU temperature, utilization, power, SM clock, memory, disk, network and LLM inference throughput in a pre-built dashboard. 一条命令搭建 DGX Spark 集群监控
Production LLM serving infrastructure using Triton Inference Server, vLLM, and Ray Serve with OpenAI-compatible endpoints. Includes Kubernetes autoscaling configs driven by DCGM GPU metrics and a BentoML packaging path for portable model deployment.
Prometheus exporter for hardware telemetry from DMTF Redfish-capable BMCs. Multi-target probe pattern. Demo stack included.
Find out how many GPU-hours you're paying for and not using
Mock NVIDIA dcgm-exporter with MIG metrics simulation for development and testing
Open-source GPU dynamic power management for datacenter — Python brain, Rust agent, Prometheus/Grafana
Freelens extension: per-pod GPU usage (utilisation, VRAM, power) from dcgm-exporter or any per-process GPU exporter, scraped via the apiserver pod-proxy. Zero cluster footprint. Companion to kubectl-gpugo.
To associate your repository with the dcgm topic, visit your repo's landing page and select "manage topics."