SessionEngineering trackconfirmed

Reinforcement Learning without Verifiable Rewards

Day
Day 3 — Session Day 2 · Wed, Jul 1
Time
1:30pm-1:50pm
Room
Track 9 · Room 2016
Track
Posttraining & Midtraining
Share
Track theme
Post-Training & Mid-Training (RL)

Post- and mid-training — what's next after RLHF, decentralized RL training at scale, and the infrastructure behind frontier models.

Accessible with the Engineering pass and above.

About this session

Verifiable rewards are the gold standard for RL training, but real-world agent tasks frequently lack clean deterministic evaluation objectives. This talk surveys our efforts to scale RL in non-verifiable settings -- including task synthesis, unsupervised environment design, and automatic judge calibration -- to ultimately enable self-improvement in production, grounded in real-world agent traces and domain-specific context.

Speaker