Evaluating agents — long-horizon benchmarks, autonomous debuggers, model whisperers, and evals-driven development.
Accessible with the Engineering pass and above.
(updated) an off-the-record session with Mihnea Munteanu, Senior Product Lead, Ask YouTube / AI Search @ Google