Data quality for frontier models and reinforcement-learning environments — curation, long-horizon training, and reliability.
Accessible with the Leadership (All-Access) pass and above.
Model releases are heralded by a flourish of trumpets, a chorus of weeping angels, and often, inflated benchmark claims. Why do benchmarks so often not reflect real-world value? Is it intrinsic to the science of benchmarking, or just the consequence of our current practices? Is LM Arena a cancer on AI?