Amazon SageMaker HyperPod and Qumulo allow training compute in one AWS Region while keeping data in another, reducing cross-Region latency and transfer costs. This solution helps teams maintain high throughput without replicating data or modifying code.
The system achieved 99 percent GPU utilization and p5.48xlarge network saturation with sub-3 ms data operations, matching the performance of co-located clusters. A HyperPod cluster in the US West (Oregon) Region matched the throughput of a cluster in the US East (Ohio) Region after a short warmup period.
The solution pairs Amazon SageMaker HyperPod with Qumulo’s Cloud Data Fabric (CDF) and Cloud Native Qumulo (CNQ). CDF uses predictive caching through NeuralCache to serve data from local NVMe storage, making cross-Region latency transparent after an initial warmup.
"We validated the approach by running the same training job independently on two clusters," said the source. "The spoke cluster matched the hub’s performance after a short warmup, demonstrating the effectiveness of the solution."
The announcement follows the release of Qumulo’s Cloud Data Fabric, which enables seamless data access across Regions. This development addresses the challenge of balancing data locality with compute flexibility in AI training.
Amazon SageMaker HyperPod and Qumulo did not specify exact pricing or availability dates, and the solution’s effectiveness depends on data access patterns and network conditions. The source noted that throughput converges within the first 100–150 batches.
Source: awsml