Amazon announced Amazon SageMaker HyperPod, a new AI service that enables multiple teams to share GPU clusters while maintaining isolation and fairness. It is the company's first major update to its SageMaker platform since its initial release in 2017.

Amazon reported that SageMaker HyperPod provides resilient, optimized clusters orchestrated by Amazon Elastic Kubernetes Service (Amazon EKS) or Slurm, allowing organizations to run distributed training, interactive development, and model inference at scale. This capability compares with its predecessor, which lacked such advanced orchestration and isolation features.

SageMaker HyperPod is built on Amazon EKS and targets organizations with multiple teams needing shared access to expensive GPU clusters for generative AI operations. Availability begins with a reference architecture for multi-tenant environments, initially for enterprises with complex team collaboration needs.

"HyperPod simplifies the management of large-scale compute clusters for gen AI workloads," said Anurag Sharma, a machine learning expert at Amazon. "It automatically handles node health monitoring, fault recovery, and cluster lifecycle management, ensuring operational efficiency for multi-team environments.

The announcement follows Amazon's focus on improving collaboration and resource efficiency in AI development. Amazon said the new service addresses the growing need for shared GPU infrastructure without compromising security or cost tracking.

Amazon did not specify the exact number of teams that can use the service simultaneously, and it raised the question of how to handle resource contention in high-demand scenarios. The company said it will continue to refine the service based on user feedback.

Source: awsml