SponsorEngineering trackconfirmed

Don't Ship Skills Without Evals

Day
Day 3 — Session Day 2 · Wed, Jul 1
Time
3:20pm-3:40pm
Room
Track 5 · Room 2005
Track
Evals
Share
Track theme
Evals

Evaluating agents — long-horizon benchmarks, autonomous debuggers, model whisperers, and evals-driven development.

Accessible with the Engineering pass and above.

About this session

There are thousands agent skills. Almost none of them are tested. They get vibe-checked with two manual runs, maybe a thumbs-up from a colleague, then shipped. You wouldn't merge code without tests — so why are we shipping skills without evals? This talk covers the full lifecycle of building reliable agent skills: what a skill actually is (and isn't), how to write one that triggers correctly, and how to build a lightweight eval harness that catches failures before your users do.

Speaker