Data quality for frontier models and reinforcement-learning environments — curation, long-horizon training, and reliability.
Accessible with the Engineering pass and above.
As autonomous agents push towards longer-horizon tasks, a number of challenges emerge in measuring and improving frontier model capabilities. In this talk, we discuss how long-horizon tasks are defined and measured, how RL environments and verifiers have to scale for more complex and open-ended tasks, and how we navigate these problems at Theta.