Reference Architectures
CloudFormation templates and deployment guides for distributed training infrastructure.
๐ฅ๏ธ
Compute
๐ฅ๏ธ ๐ง โธ๏ธ ๐ฆ ๐งฎ
SageMaker HyperPod
Managed GPU clusters with resilient training and automatic recovery
AWS ParallelCluster
HPC cluster management with Slurm scheduler for distributed training
Amazon EKS
Kubernetes-based orchestration for distributed training jobs
AWS Batch
Serverless batch computing for distributed training workloads
AWS Parallel Computing Service
Managed Slurm service for HPC and ML workloads
๐
Networking
๐พ
Storage
๐ก๏ธ