SessionLeadership trackconfirmed

When Will The Benchmaxxing Plague End?

Day
Day 2 — Session Day 1 · Tue, Jun 30
Time
2:50pm-3:10pm
Room
Track 9 · Room 2016
Track
AI Architects: Show my Workflow
Share
Track theme
Data Quality & RL Environments

Data quality for frontier models and reinforcement-learning environments — curation, long-horizon training, and reliability.

Accessible with the Leadership (All-Access) pass and above.

About this session

Model releases are heralded by a flourish of trumpets, a chorus of weeping angels, and often, inflated benchmark claims. Why do benchmarks so often not reflect real-world value? Is it intrinsic to the science of benchmarking, or just the consequence of our current practices? Is LM Arena a cancer on AI?

Speaker