Training Frameworks
Browse test cases and training examples organized by framework and library.
PyTorch
22 test casesPyTorch FSDP
Fully Sharded Data Parallel training for large language models
PyTorch DDP
Distributed Data Parallel training - the foundation for multi-GPU PyTorch
DeepSpeed
Microsoft DeepSpeed ZeRO optimizer for memory-efficient distributed training
TorchTitan
PyTorch native distributed training framework for production LLM pre-training
Picotron
Lightweight distributed training library for educational and research use
vLLM
High-throughput LLM inference and serving engine
OpenRLHF
Open-source RLHF framework for training reward models and policy optimization
NVIDIA Dynamo
Distributed LLM inference with KV cache-aware routing and disaggregated prefill/decode on HyperPod EKS
MosaicML Composer
Training efficiency library with algorithmic speedups and multi-GPU orchestration
NVIDIA Isaac Lab
Sim-to-real robot learning with NVIDIA Isaac Lab on GPU clusters
OpenVLA OFT
Open Vision-Language-Action models with fine-tuning for robotic manipulation
nanoVLM
Lightweight vision-language model training for embodied AI
V-JEPA 2
Video Joint Embedding Predictive Architecture for physical world understanding
Cosmos 3
NVIDIA Cosmos 3 Physical AI flywheel โ omnimodal world models for generate โ post-train โ eval
DreamZero
14B World-Action Model (WAM) for robotic manipulation via video diffusion on EKS
V-JEPA 2.1
Updated Video Joint Embedding Predictive Architecture for physical world understanding
PointWorld
Distributed 3D world model pre-training for robotic manipulation (NVIDIA + Stanford)
OpenVLA
Open Vision-Language-Action model for generalist robotic manipulation
TRL (Transformers Reinforcement Learning)
HuggingFace TRL for RLHF, DPO, PPO, and reward model training
vERL
Scalable reinforcement learning framework for LLM alignment and post-training
SLIME
Lightweight distributed training library for efficient LLM fine-tuning
Model Distillation
Knowledge distillation for compressing large models into smaller, efficient ones
Megatron / NeMo
6 test casesMegatron-LM
NVIDIA's framework for training multi-billion parameter transformer models
NVIDIA NeMo
End-to-end framework for building, training, and deploying AI models
NeMo RL
Reinforcement learning from human feedback with NeMo
BioNeMo
NVIDIA's framework for biomolecular AI model training
Megatron-Bridge
NVIDIA Megatron-Bridge + UCCL-EP for MoE training with expert-parallel all-to-all over EFA
NeMo 1.0 (Legacy)
Legacy NeMo 1.0 training examples โ superseded by NeMo 2.x
JAX
1 test casesAWS Neuron
2 test casesReinforcement Learning
4 test casesNVIDIA Isaac Lab
Sim-to-real robot learning with NVIDIA Isaac Lab on GPU clusters
TRL (Transformers Reinforcement Learning)
HuggingFace TRL for RLHF, DPO, PPO, and reward model training
vERL
Scalable reinforcement learning framework for LLM alignment and post-training
SLIME
Lightweight distributed training library for efficient LLM fine-tuning
Model Customisation
1 test casesPhysical AI & Robotics
9 test casesNVIDIA Isaac Lab
Sim-to-real robot learning with NVIDIA Isaac Lab on GPU clusters
OpenVLA OFT
Open Vision-Language-Action models with fine-tuning for robotic manipulation
nanoVLM
Lightweight vision-language model training for embodied AI
V-JEPA 2
Video Joint Embedding Predictive Architecture for physical world understanding
Cosmos 3
NVIDIA Cosmos 3 Physical AI flywheel โ omnimodal world models for generate โ post-train โ eval
DreamZero
14B World-Action Model (WAM) for robotic manipulation via video diffusion on EKS
V-JEPA 2.1
Updated Video Joint Embedding Predictive Architecture for physical world understanding
PointWorld
Distributed 3D world model pre-training for robotic manipulation (NVIDIA + Stanford)
OpenVLA
Open Vision-Language-Action model for generalist robotic manipulation