V-JEPA 2.1

Updated Video Joint Embedding Predictive Architecture for physical world understanding

View on GitHub

Overview

V-JEPA 2.1 is an updated version of V-JEPA 2 (Video Joint Embedding Predictive Architecture) that learns representations of the physical world from video through self-supervised learning. It predicts future states without pixel-level reconstruction.

Key Features

  • Self-supervised video representation learning
  • Physical world understanding from unlabeled video
  • Improved training stability over V-JEPA 2
  • Scales to large video datasets
  • Transfer to robotics and embodied AI tasks