
VP of Kernels at Together AI and Assistant Professor of Computer Science and Engineering at UC San Diego, focused on efficient machine learning systems and GPU performance.
Using publicly available information we constructed an analysis to help you get a feel for this speaker before deciding to attend their session.
Fu leads Together AI's GPU-kernels team and is a lead author on ThunderKittens, FlashFFTConv, and Monarch Mixer, so his session is a chance to hear from a hardware-aware-ML researcher how to make models actually fast on modern GPUs — and why he argues software-hardware co-design (not just bigger models) unlocks the next order of magnitude of performance.