๐ฏ
Reinforcement Learning
4 test casesRLHF, DPO, PPO, and scalable RL frameworks for LLM alignment and post-training. Train reward models, optimize policies, and align models with human preferences at scale.
๐ค
NVIDIA Isaac Lab
Sim-to-real robot learning with NVIDIA Isaac Lab on GPU clusters
Isaac LabRoboticsSim2RealPhysical AI
๐ฏ
TRL (Transformers Reinforcement Learning)
HuggingFace TRL for RLHF, DPO, PPO, and reward model training
TRLRLHFDPOPPOAlignment
๐ฏ
vERL
Scalable reinforcement learning framework for LLM alignment and post-training
vERLRLHFPPOScalable RLAlignment
๐ฏ
SLIME
Lightweight distributed training library for efficient LLM fine-tuning
SLIMELightweightFine-tuningEfficient