New in Together GPU Clusters: Reliability and control for production GPU clusters

Together AI detailed a set of changes to its Together GPU Clusters platform, grouped into platform health and operational control. The company said the work targets running training and inference at scale.
On platform health, Together AI introduced passive health checks that run continuously across every node, observing real workloads, logs and metrics. The stated coverage list includes GPUs falling off the bus, thermal throttling, Xid errors, Slurm node drains on failure, and an expanding set of hardware and software failure signals. Together AI said the checks observe live workloads with near-zero overhead on running jobs.
Passive checks were paired with auto node repair. When monitoring detects a node-level issue, it generates a recommended remediation for an operator to review. Together AI described four repair actions mapped automatically based on the failure signature: reboot, reprovision, failover, and remove. The company described the approach as human-in-the-loop, with the system detecting and recommending, the customer approving, and Together handling recovery, and said it is building fully automated node-repair options for specific failure modes.
Together AI also rebuilt its Slurm-on-Kubernetes stack, based on its fork of the open-source project Slinky, under the name Together Slurm-on-K8s 2.0. Stated changes include self-healing worker daemons, automatic reaping of orphan processes, job accounting moved to PVC-backed storage, kernel-level tracking and cleanup of a job's descendants, and rebuilding Slurm's GPU view on every node start. The stack exposes DCGM metrics in cluster Grafana dashboards. Together AI said newly provisioned Slurm clusters run on the latest stack by default, and existing managed Slurm clusters can be migrated in place at no cost via a maintenance window.
For operational control, Together AI added a rebuilt cluster details view showing node health states (healthy, booting, unhealthy, pending, paused), live usage metrics across GPU nodes including utilization, memory and network bandwidth, drill-down to Grafana, an event timeline, and cluster configuration. Three tabs cover nodes, health checks, and repair.
The company added external OIDC support for Kubernetes RBAC, letting clusters authenticate against an existing identity provider including Google, Okta, Auth0, Microsoft Entra ID or another OIDC-compatible provider, with per-user kubectl access and standard Kubernetes RBAC via ClusterRoleBindings and RoleBindings. Together AI said external OIDC must be configured at cluster creation time, and that OIDC support for Slurm and Kubernetes clusters is coming soon.
Startup scripts let customers customize Slurm worker nodes, login nodes and the controller via shell scripts that fire at node boot, job start and job end. Scripts are configured in the Together Cloud console, can be applied to new clusters at creation time, and are validated for errors.
Together AI also described acceptance testing as an opt-in for larger and longer-running clusters. Acceptance tests, which validate GPU health, networking and storage, are skipped by default for smaller or short-lived clusters in favor of faster availability; Together AI recommends enabling them at cluster creation for multi-GPU training or long-running production clusters.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.