SessionExpo trackconfirmed

Optimizing Open Models for Production Grade Inference

Day
Day 4 — Session Day 3 · Thu, Jul 2
Time
2:25pm-2:45pm
Room
Expo Stage 1 NE
Track
Share
Track theme
Harness Engineering & Agent Memory (Expo)

Sponsor talks on harness engineering as a core skill, agent memory vs. learning, verifiers, and relational context engines.

Accessible with the Expo Explorer pass and above.

About this session

Open-source foundation models are rapidly closing the gap with proprietary systems, enabling organizations to build powerful AI applications with greater flexibility and control. However, deploying these models in production introduces a new set of challenges: latency, throughput, scalability, and cost efficiency.In this talk, we'll explore the modern inference optimization techniques that power large-scale AI systems in production. Topics include KV cache optimization, cache-aware routing, prefill/decode disaggregation, speculative decoding, and other emerging approaches used to improve performance and reduce infrastructure costs.Through practical examples and real-world architecture patterns, attendees will gain a deeper understanding of how to run open models efficiently at scale.

Speakers