Viewpoint

Sparse experts can combine large capacity with computational efficiency

Sparse experts can combine large capacity with computational efficiency

Dean says mixture-of-experts architectures allow models to maintain very large capacity while activating only the components most useful for a particular token or request, improving efficiency over dense models.

Speaker
Jeff Dean
Source timestamp
7:00

More from this interview