← Catalog
PaperTsinghua University · 2024
iVideoGPT: Interactive VideoGPTs are Scalable World Models
A scalable autoregressive transformer world model with a compressive tokenizer, pretrained on large video corpora for interactive prediction and control.
Summary
iVideoGPT casts world modeling as interactive video prediction: a compressive tokenizer plus an autoregressive transformer predict future frames conditioned on actions, and the model scales via pretraining on millions of trajectories. Fine-tunes to action-conditioned prediction, model-based RL, and visual planning.
Metadata
- Authors
- Wu et al., Tsinghua
- Year
- 2024
- arXiv
- 2405.15223
- Venue
- NeurIPS 2024
Relationships
Introduces
iVideoGPT
#autoregressive#video#model-based-rl