New in Together GPU Clusters: Reliability and control for production GPU clusters
Together AI
Together AI released updates to its GPU Clusters platform including passive health checks that monitor running workloads for failures like GPU bus drops and thermal throttling, auto node repair with human approval, and a rebuilt Slurm-on-Kubernetes stack addressing daemon crashes and process cleanup. The platform added operational features including a redesigned cluster overview dashboard showing health and utilization, external OIDC authentication for per-user Kubernetes access, and startup scripts for self-serve node customization. These changes reduce incident resolution time from hours to minutes and enable teams to manage clusters at scale without sharing admin credentials or performing manual node setup.
Why it matters
See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.