Viewpoint
Waymo's foundation model is a multimodal world-action language model using encoder-decoder and fast/slow reasoning.
Waymo's foundation model is a multimodal world-action language model using encoder-decoder and fast/slow reasoning.
The Waymo Foundation model is a multimodal, world-action language model with an encoder-decoder architecture and fast and slow reasoning paths to handle both millisecond safety decisions and complex semantic understanding.
- Speaker
- Dmitri Dolgov
- Source timestamp
- 24:41
More from this interview
- Physical AI requires moving fast while shipping safely
- Physical AI faces life-cost errors, extreme latency, no digitized internet, and pre-deployment safety validation.
- An 18-month demo versus a 15-year scalable service with over 200 million miles.
- Reliability follows an exponential ladder of nines; at fleet scale rare events become daily.
- Complementary camera, LiDAR, and radar with active sensors provide redundant, superhuman safety.
- Do not anchor to today's component costs
- Waymo rebuilt its stack around AI breakthroughs; production without regressions is harder than prototyping.
- Structured approaches should channel scale, not fight it
- Waymo's structure-augmented models improve safety validation
- Waymo's simulator is a large AI model for closed-loop evaluation and training.
- Generative models enable training on rare edge-case scenarios
- Agent, simulator, and critic AIs share a foundation world model and data flywheel.
- Metrics and evaluation build the strategic moat
- Autonomy trust is earned through safety frameworks and data
- Physical AI mirrors digital AI's earlier stage; next decade unfolds in physical world.