Viewpoint
Jointly trained world-action models are faster for test-time planning.
Jointly trained world-action models are faster for test-time planning.
The speaker argues that jointly training a single model to output both actions and next states is more efficient and faster for planning than using separate, sequential models.
- Speaker
- Francois Chaubard
- Topic
- World Models
- Source timestamp
- 18:25
More from this interview
- Intelligence per sample is a major unsolved problem in AI.
- Perfect sample efficiency is achievable with a perfect world model.
- World models enable pre-programmed planning without environmental sampling.
- Control problems are solved using transition functions and policies.
- Non-differentiable stochastic systems require reinforcement learning methods.
- World models help estimate value functions and joint state-action distributions.
- World models and policies are combined into a joint probability distribution.
- Action conditioning is added to world models to enable interaction.
- AlphaGo uses Monte Carlo Tree Search for expensive test-time planning.
- MCTS scaling is limited by large action spaces and real-time requirements.
- Self-driving car state and action spaces are effectively infinite.
- Self-driving requires modeling the impact of actions on other agents.
- Tesla's fleet provides a unique dataset of driver actions.
- Model-free reinforcement learning predicts actions directly from states.
- Model-based reinforcement learning uses a world model for planning.
- Dreamer trains policies on synthetic data from a world model.
- Video generation models can be fine-tuned into actionable world models.
- Dreamer V4 for robotics uses pre-trained diffusion models and tele-operation data.
- JEPA operates in latent space to improve sample efficiency.
- A major open problem is achieving high-fidelity predictions in world models.
- Real-time adaptation and estimation of changing dynamics remain challenging.