Yunmo Koo is a founding engineer at FriendliAI focused on LLM inference optimization, distributed training, multi-cloud systems, and LLMOps. He builds production ML infrastructure for lower latency, better reliability, and improved cost efficiency.
Using publicly available information we constructed an analysis to help you get a feel for this speaker before deciding to attend their session.
Yunmo Koo is a Founding Engineer at FriendliAI who has spent over five years building the core LLM inference runtime — speculative decoding, KV-cache optimization, GPU kernels, and continuous batching — and now applies that deep systems knowledge to the demands of agentic AI. His session offers first-hand insight into what it takes to make multi-step agent workloads fast and cost-efficient in production.