Viewpoint

Video generation models can be fine-tuned into actionable world models.

Video generation models can be fine-tuned into actionable world models.

The speaker explains that state-of-the-art video diffusion models (like Sora) can be fine-tuned with a small amount of action-conditioned data to create effective world models for robotics and self-driving.

Source timestamp
51:29

More from this interview