SessionEngineering trackconfirmed

Large clusters for small models

Day
Day 4 — Session Day 3 · Thu, Jul 2
Time
1:55pm-2:15pm
Room
Track 9 · Room 2016
Track
Inference
Share
Track theme
Inference Engineering

Inference engineering — distributed inference, routing, KV-cache-aware scheduling, ASICs, and deep vLLM debugging.

Accessible with the Engineering pass and above.

About this session

Small task-specific models are cheaper, faster and narrowly better than the frontier. But a wide catalog of small models is tricky to serve - dedicated worker pools sit idle, top-down request routers choke up on the huge volume of small requests, your users bring 100s of LoRAs.. In this talk we show how we serve 1M tokens per second with small models, how we architect our cluster for maximum throughput AND minimum latency and how we apply autoresearch to rewrite our inference code to support 10+ new models a week.

Speaker