V-JEPA 2.1
Updated Video Joint Embedding Predictive Architecture for physical world understanding
View on GitHubOverview
V-JEPA 2.1 is an updated version of V-JEPA 2 (Video Joint Embedding Predictive Architecture) that learns representations of the physical world from video through self-supervised learning. It predicts future states without pixel-level reconstruction.
Key Features
- Self-supervised video representation learning
- Physical world understanding from unlabeled video
- Improved training stability over V-JEPA 2
- Scales to large video datasets
- Transfer to robotics and embodied AI tasks