worldmodels.fyi
← Catalog
PaperMeta AI (FAIR) · 2025

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Action-free joint-embedding predictive video model pretrained on 1M+ hours of video, post-trained for robot planning.

Summary

Trains a self-supervised video model (V-JEPA 2) on over one million hours of internet video with a joint-embedding predictive objective, reaching strong motion understanding and action anticipation. A latent action-conditioned variant (V-JEPA 2-AC) enables zero-shot robot planning on Franka arms from image goals.

Metadata

Authors
Meta AI (FAIR)
Year
2025
arXiv
2506.09985
Introduces
V-JEPA 2

Relationships

Introduces

Uses

DROID (robot data)

Evaluated on

Something-Something v2
#self-supervised#jepa#planning