
Researcher at MiniMax focused on reinforcement learning and model evaluation for the M-series models.
Using publicly available information we constructed an analysis to help you get a feel for this speaker before deciding to attend their session.
Olive Song (Jiayuan Song) is a senior RL researcher at MiniMax who leads reinforcement learning training behind the company's M-series open-weight models. She is a co-author of the MiniMax-M1 paper and the CISPO algorithm, and speaks candidly about hard-won production RL lessons — debugging FP32 precision in the LM head as a convergence blocker, fighting reward hacking, and building perturbation pipelines for robust agentic models. Her session offers practitioner-grade insight from someone shipping frontier open models at scale.