Post- and mid-training — what's next after RLHF, decentralized RL training at scale, and the infrastructure behind frontier models.
Accessible with the Engineering pass and above.
Verifiable rewards are the gold standard for RL training, but real-world agent tasks frequently lack clean deterministic evaluation objectives. This talk surveys our efforts to scale RL in non-verifiable settings -- including task synthesis, unsupervised environment design, and automatic judge calibration -- to ultimately enable self-improvement in production, grounded in real-world agent traces and domain-specific context.