Viewpoint
Non-differentiable stochastic systems require reinforcement learning methods.
Non-differentiable stochastic systems require reinforcement learning methods.
When an environment becomes stochastic and non-differentiable (e.g., due to an unpredictable adversary), direct optimization fails, necessitating reinforcement learning techniques like value iteration or actor-critic.
- Speaker
- Francois Chaubard
- Source timestamp
- 12:41
More from this interview
- Intelligence per sample is a major unsolved problem in AI.
- Perfect sample efficiency is achievable with a perfect world model.
- World models enable pre-programmed planning without environmental sampling.
- Control problems are solved using transition functions and policies.
- World models help estimate value functions and joint state-action distributions.
- World models and policies are combined into a joint probability distribution.
- Action conditioning is added to world models to enable interaction.
- Jointly trained world-action models are faster for test-time planning.
- AlphaGo uses Monte Carlo Tree Search for expensive test-time planning.
- MCTS scaling is limited by large action spaces and real-time requirements.
- Self-driving car state and action spaces are effectively infinite.
- Self-driving requires modeling the impact of actions on other agents.
- Tesla's fleet provides a unique dataset of driver actions.
- Model-free reinforcement learning predicts actions directly from states.
- Model-based reinforcement learning uses a world model for planning.
- Dreamer trains policies on synthetic data from a world model.
- Video generation models can be fine-tuned into actionable world models.
- Dreamer V4 for robotics uses pre-trained diffusion models and tele-operation data.
- JEPA operates in latent space to improve sample efficiency.
- A major open problem is achieving high-fidelity predictions in world models.
- Real-time adaptation and estimation of changing dynamics remain challenging.