Share GPU clusters across teams with isolation and fairness using Amazon SageMaker HyperPod

AWS has published a reference architecture for running a shared, multi-tenant Amazon SageMaker HyperPod cluster orchestrated by Amazon Elastic Kubernetes Service, aimed at companies whose teams must share costly GPU capacity while keeping isolation, resource fairness and operational independence. Amazon SageMaker HyperPod is described as an AI service that manages large-scale compute clusters and automatically handles node health monitoring, fault recovery and cluster lifecycle management.
The blueprint connects user identity to isolated workloads through layered controls. Authentication runs through AWS IAM Identity Center, which federates with an external identity provider; the example uses Microsoft Entra ID with groups named TeamA, TeamB and Admin, synchronized via SCIM and authenticated with SAML 2.0. Two access paths exist: the CLI, where users run aws sso login to get temporary credentials from a team permission set and submit tasks with kubectl, and the Identity Center portal, where users select SageMaker Studio and enter a team-specific SageMaker AI domain.
Authorization splits across two levels. Each team gets a dedicated IAM role serving as its SageMaker AI domain execution role, with policies for SageMaker AI, scoped Amazon S3 prefixes, CloudWatch and the EKS permissions eks:AccessKubernetesApi and eks:MutateViaKubernetesApi. The trust policy must include sagemaker.amazonaws.com, and pods.eks.amazonaws.com if the role is reused as an EKS Pod Identity association. On the cluster boundary, EKS access entries map those IAM roles to Kubernetes RBAC policies scoped to the team namespace, so requests from Studio or kubectl can only reach that team's resources. Separate Identity Center permission sets, such as TeamA-permission-set and TeamB-permission-set, define CLI permissions independently of the Studio role.
Inside the HyperPod EKS cluster, cross-cutting layers handle HyperPod Observability and HyperPod Task Governance, the latter for compute quota management and scheduling priorities. Namespace A and Namespace B host each team's Spaces for interactive development, PyTorch training jobs and inference endpoints. Storage uses POSIX-compliant Amazon FSx for Lustre or Amazon FSx for OpenZFS with per-team directories such as /fsx/TeamA and per-user homes, plus per-team or shared Amazon S3 buckets governed by the team's IAM execution role.
Why it matters: Organizations adopting this pattern can attribute shared GPU spend per team through namespace-level cost allocation and chargeback, while teams keep an isolated operating boundary from authentication through workload execution.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
A reference architecture for securely sharing one Amazon SageMaker HyperPod EKS cluster across multiple teams, using AWS IAM Identity Center for authentication, per-team SageMaker Domains and Kubernetes namespaces for isolation, HyperPod Task Governance for fairness, and namespace-level cost allocation for chargeback.