
Simran Arora is a Stanford computer science PhD student working on AI systems, GPU kernels, and efficient model execution. She is a co-author of KernelBench and related multi-GPU kernel research, and is associated with Together AI in recent AI-for-science and systems work.
Using publicly available information we constructed an analysis to help you get a feel for this speaker before deciding to attend their session.
Simran Arora is a Stanford PhD researcher (advised by Chris Ré) working at the intersection of ML systems and model design. She is a core author/lead of hardware-efficient AI projects: ThunderKittens (a CUDA tile-primitive DSL for fast GPU kernels, ~3.5k stars), the BASED attention-free / linear-attention architecture, and KernelBench (an LLM-GPU-kernel benchmark). Her work pushes the Pareto frontier of model quality vs. compute efficiency, with linear-attention ideas adopted across newer architectures (Mamba-v2, RWKV-v5/v6, MiniMax).