← Catalog
PaperMeta AI (FAIR) · 2025
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Action-free joint-embedding predictive video model pretrained on 1M+ hours of video, post-trained for robot planning.
Summary
Trains a self-supervised video model (V-JEPA 2) on over one million hours of internet video with a joint-embedding predictive objective, reaching strong motion understanding and action anticipation. A latent action-conditioned variant (V-JEPA 2-AC) enables zero-shot robot planning on Franka arms from image goals.
Metadata
- Authors
- Meta AI (FAIR)
- Year
- 2025
- arXiv
- 2506.09985
- Introduces
- V-JEPA 2
Relationships
#self-supervised#jepa#planning