Interview

Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club

At our latest YC Paper Club, researchers and builders presented on multi-GPU kernel optimization, intelligence per watt for local inference, AI-generated GPU kernels and benchmarking, heterogeneous inference infrastructure design, and GPU-accelerated game engines for reinforcement learning. Thanks to the following presenters:

Stuart Sul (Stanford / Cursor), John (Stanford), Mark (PyTorch / GPU Mode / CoreAuto), Misha (Marlo), and Brennan (Stanford)


Chapters:

0:00 – Francois Chaubard: The case for chip and kernel specialization

7:16 – Stuart Sul: Parallel Kittens - Systematic and Practical Simplification of Multi-GPU Al Kernels (https://arxiv.org/abs/2511.13940)

21:29 – Jon Saad-Falcon: Intelligence per Watt - Measuring the Intelligence Efficiency of Local and Cloud AI (https://arxiv.org/abs/2511.07885)

31:05 – Mark Saroufim: When Al Starts Writing Systems Code

47:04 – Misha Smelyanskiy: Why AI Inference Needs Heterogeneous Hardware

1:04:33 – Brennan Shacklett: Building a High-Throughput Game Engine that Runs ENTIRELY on the GPU (https://madrona-engine.github.io/shac...)

1:15:35 – Wrap-up & what's next

Source attribution: Original source

Report issue

Interview details

Publication date
Aug 15, 2026
Duration
1h 16m

Speakers

Viewpoints in this interview