OpenVLA

Open Vision-Language-Action model for generalist robotic manipulation

View on GitHub

Overview

OpenVLA (Open Vision-Language-Action) is a generalist robotic manipulation model that combines vision understanding, language grounding, and action prediction in a single architecture. It can be used as a foundation policy for diverse manipulation tasks.

Key Features

  • Generalist policy โ€” Single model handles diverse manipulation tasks
  • Vision-Language-Action โ€” Integrates visual perception, language understanding, and motor control
  • Open source โ€” Openly available model weights and training code
  • Transfer learning โ€” Pre-trained on diverse robot datasets for broad generalization

See also OpenVLA-OFT for the Orthogonal Fine-Tuning variant optimized for task-specific adaptation.