Intelligence per sample is a major unsolved problem in AI.The speaker defines the core challenge as measuring how much smarter a model becomes with each additional data point, noting that current models are very inefficient at learning from small datasets compared to humans.Francois ChaubardSample EfficiencySource at 1:01Save viewpoint
Perfect sample efficiency is achievable with a perfect world model.The speaker explains that zero-sample learning is theoretically possible if one has a flawless predictive model of the environment, using Newtonian physics and model predictive control as an example of applying known laws to plan actions without real-world trials.Francois ChaubardWorld ModelsSource at 3:00Save viewpoint
World models enable pre-programmed planning without environmental sampling.Using the example of NASA's asteroid interception, the speaker illustrates that a perfect world model allows for pre-calculated trajectories that don't require adaptive sampling from the environment during execution.Ankit GuptaWorld ModelsSource at 3:37Save viewpoint
Control problems are solved using transition functions and policies.The speaker defines the core elements of a control problem: a state, an action (control input), a world model (state transition function), and a policy (what action to take given a state).Francois ChaubardReinforcement LearningSource at 8:26Save viewpoint
Non-differentiable stochastic systems require reinforcement learning methods.When an environment becomes stochastic and non-differentiable (e.g., due to an unpredictable adversary), direct optimization fails, necessitating reinforcement learning techniques like value iteration or actor-critic.Francois ChaubardReinforcement LearningSource at 12:41Save viewpoint
World models help estimate value functions and joint state-action distributions.The speaker clarifies that in standard reinforcement learning, machine learning models are used to estimate value functions, and world models provide a framework for jointly modeling state transitions and policies to create more intelligent behavior.Ankit GuptaReinforcement LearningSource at 16:19Save viewpoint
World models and policies are combined into a joint probability distribution.The speaker describes the desired joint distribution for intelligent agents, which factorizes into a policy and a world model, and notes these are often learned separately before being combined.Francois ChaubardWorld ModelsSource at 16:46Save viewpoint
Action conditioning is added to world models to enable interaction.Building on passive video prediction, action conditioning is injected into world models to allow them to simulate the consequences of actions, a technique central to the Dreamer paper series.Francois ChaubardWorld ModelsSource at 17:36Save viewpoint
Jointly trained world-action models are faster for test-time planning.The speaker argues that jointly training a single model to output both actions and next states is more efficient and faster for planning than using separate, sequential models.Francois ChaubardWorld ModelsSource at 18:25Save viewpoint
AlphaGo uses Monte Carlo Tree Search for expensive test-time planning.The speaker details how AlphaGo uses MCTS, performing many model invocations to build a search tree that balances exploration and exploitation to select optimal moves in games with small action spaces.Francois ChaubardMonte Carlo Tree SearchSource at 26:40Save viewpoint
MCTS scaling is limited by large action spaces and real-time requirements.The speaker explains that the MCTS approach used in AlphaGo does not scale to domains like self-driving or robotics because of large action spaces, non-deterministic environments, and the need for real-time decisions.Francois ChaubardMonte Carlo Tree SearchSource at 33:14Save viewpoint
Self-driving car state and action spaces are effectively infinite.The speaker describes the state space of a self-driving car as massive (including surroundings, vehicle state, weather) and calculates that even a simplified action space (steering, brake, gas) is vastly larger than in games like Go.Ankit GuptaSelf-DrivingSource at 34:14Save viewpoint
Self-driving requires modeling the impact of actions on other agents.The speaker notes that while Newtonian physics can model car dynamics, the non-differentiable element is predicting how other road users will react to one's own actions, which is essential for safe navigation.Francois ChaubardSelf-DrivingSource at 36:59Save viewpoint
Tesla's fleet provides a unique dataset of driver actions.The speaker highlights a key competitive advantage for Tesla: its fleet of consumer vehicles collects video data paired with the driver's actual actions (steering, braking), which most other companies lack.Francois ChaubardSelf-DrivingSource at 39:31Save viewpoint
Model-free reinforcement learning predicts actions directly from states.The speaker defines model-free reinforcement learning as a method that learns a policy (action from state) without an explicit model of the environment's dynamics, analogous to behavior cloning or next-token prediction.Francois ChaubardReinforcement LearningSource at 41:28Save viewpoint
Model-based reinforcement learning uses a world model for planning.The speaker contrasts model-free methods with model-based RL, where an explicit world model (the transition function) is used to plan actions, enabling stronger policies but requiring more computation for test-time inference.Francois ChaubardReinforcement LearningSource at 42:21Save viewpoint
Dreamer trains policies on synthetic data from a world model.The speaker describes the Dreamer series as pioneering work that trains a world model from environment data and then generates synthetic rollouts to train a policy, achieving results like mining diamonds in Minecraft using only synthetic data.Francois ChaubardDreamer ArchitectureSource at 48:55Save viewpoint
Video generation models can be fine-tuned into actionable world models.The speaker explains that state-of-the-art video diffusion models (like Sora) can be fine-tuned with a small amount of action-conditioned data to create effective world models for robotics and self-driving.Francois ChaubardDreamer ArchitectureSource at 51:29Save viewpoint
Dreamer V4 for robotics uses pre-trained diffusion models and tele-operation data.The speaker cites a paper applying the Dreamer concept to robotics, which starts with an open-source video diffusion model and fine-tunes it with about 500 hours of tele-operation data to achieve cross-embodiment tasks.Ankit GuptaDreamer ArchitectureSource at 53:14Save viewpoint
JEPA operates in latent space to improve sample efficiency.The speaker introduces Joint Embedding Predictive Architecture (JEPA) as a method that compresses high-dimensional state spaces (like images) into latent vectors and performs world modeling and predictions in that efficient latent space.Francois ChaubardJoint Embedding Predictive ArchitectureSource at 58:49Save viewpoint