Training Frameworks

Browse test cases and training examples organized by framework and library.

๐Ÿ”ฅ

PyTorch

22 test cases
๐Ÿ”ฅ

PyTorch FSDP

Fully Sharded Data Parallel training for large language models

FSDPShardingLarge ModelsMulti-GPU
๐Ÿ”ฅ

PyTorch DDP

Distributed Data Parallel training - the foundation for multi-GPU PyTorch

DDPData ParallelMulti-GPUBaseline
๐Ÿ”ฅ

DeepSpeed

Microsoft DeepSpeed ZeRO optimizer for memory-efficient distributed training

DeepSpeedZeROMemory EfficientLarge Models
๐Ÿ”ฅ

TorchTitan

PyTorch native distributed training framework for production LLM pre-training

TorchTitanPre-training4D ParallelismProduction
๐Ÿ”ฅ

Picotron

Lightweight distributed training library for educational and research use

PicotronLightweightEducationalResearch
๐Ÿš€

vLLM

High-throughput LLM inference and serving engine

vLLMInferenceServingPagedAttention
๐Ÿ”ฅ

OpenRLHF

Open-source RLHF framework for training reward models and policy optimization

RLHFPPODPOAlignment
๐Ÿš€

NVIDIA Dynamo

Distributed LLM inference with KV cache-aware routing and disaggregated prefill/decode on HyperPod EKS

DynamoInferenceKV CacheDisaggregatedSGLang
๐Ÿ”ฅ

MosaicML Composer

Training efficiency library with algorithmic speedups and multi-GPU orchestration

MosaicMLComposerTraining EfficiencySpeedups
๐Ÿค–

NVIDIA Isaac Lab

Sim-to-real robot learning with NVIDIA Isaac Lab on GPU clusters

Isaac LabRoboticsSim2RealPhysical AIReinforcement Learning
๐Ÿค–

OpenVLA OFT

Open Vision-Language-Action models with fine-tuning for robotic manipulation

OpenVLAVLARoboticsFine-tuningPhysical AI
๐Ÿค–

nanoVLM

Lightweight vision-language model training for embodied AI

nanoVLMVLMMultimodalPhysical AIVision-Language
๐Ÿค–

V-JEPA 2

Video Joint Embedding Predictive Architecture for physical world understanding

V-JEPA 2VideoSelf-supervisedPhysical AIWorld Models
๐ŸŒ

Cosmos 3

NVIDIA Cosmos 3 Physical AI flywheel โ€” omnimodal world models for generate โ†’ post-train โ†’ eval

CosmosWorld ModelsPhysical AIOmnimodalVideo Generation
๐ŸŒ

DreamZero

14B World-Action Model (WAM) for robotic manipulation via video diffusion on EKS

DreamZeroWorld ModelsPhysical AIRoboticsVideo Diffusion
๐ŸŒ

V-JEPA 2.1

Updated Video Joint Embedding Predictive Architecture for physical world understanding

V-JEPA 2VideoSelf-supervisedPhysical AIWorld Models
๐ŸŒ

PointWorld

Distributed 3D world model pre-training for robotic manipulation (NVIDIA + Stanford)

PointWorld3D World ModelsPhysical AIRoboticsPoint Flow
๐Ÿค–

OpenVLA

Open Vision-Language-Action model for generalist robotic manipulation

OpenVLAVLARoboticsPhysical AIVision-Language-Action
๐ŸŽฏ

TRL (Transformers Reinforcement Learning)

HuggingFace TRL for RLHF, DPO, PPO, and reward model training

TRLRLHFDPOPPOAlignmentReinforcement Learning
๐ŸŽฏ

vERL

Scalable reinforcement learning framework for LLM alignment and post-training

vERLRLHFPPOScalable RLAlignmentReinforcement Learning
๐ŸŽฏ

SLIME

Lightweight distributed training library for efficient LLM fine-tuning

SLIMELightweightFine-tuningEfficientReinforcement Learning
๐Ÿงช

Model Distillation

Knowledge distillation for compressing large models into smaller, efficient ones

DistillationKnowledge TransferCompressionModel Customisation
โšก

Megatron / NeMo

6 test cases
๐Ÿงฌ

JAX

1 test cases
๐Ÿง 

AWS Neuron

2 test cases
๐ŸŽฏ

Reinforcement Learning

4 test cases
๐Ÿงช

Model Customisation

1 test cases
๐Ÿค–

Physical AI & Robotics

9 test cases