Viewpoint
Specialized systems for inference phases improve interactivity and TCO.
Specialized systems for inference phases improve interactivity and TCO.
Inference is highly heterogeneous across phases, and co-designing specialized systems for prefill, decode, attention, and MoE kernels can improve interactivity and TCO, though full-stack integration with networking and data center design adds complexity.
- Interview
- Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club
- Speaker
- Misha Smelyanskiy
- Source timestamp
- 47:04
More from this interview
- Growing token demand drives specialization in AI chip architectures.
- ParallelKittens demonstrates efficient multi-GPU kernels with minimal code.
- Local models route most inference on-device for energy and cost savings.
- Robust benchmarks are needed to evaluate AI-generated GPU kernels.
- Batch simulators achieve 100x RL speedups but need simpler GPU programming.