
Serving New LLM Architectures in Production
Description
Modern LLM architectures are evolving fast, but turning new ideas into reliable, high-performance production serving remains a systems challenge. Join vLLM × Ant Ling for a technical conversation on what it takes to support and optimize emerging LLM architectures in practice. From runtime abstractions and memory management to model integration, profiling, and accelerator-aware optimization, we’ll explore how engineering teams move from initial architecture support to real-world production performance. Whether you build inference infrastructure, develop models, or optimize AI systems at scale, this session will offer practical insights into bringing next-generation LLM architectures into production.
Let your network know you`re going
Share this event to start conversations, invite colleagues, and connect before it begins.