
Modular at NVIDIA GTC 2026
Beschrijving
Writing high-performance GPU code is unreasonably hard, and it's getting worse. The bar for peak TFLOPS keeps rising, hardware keeps changing, and only a handful of engineers on the planet can write this code well. Modular is changing that. Meet us at NVIDIA GTC 2026 and watch Mojo and MAX push NVIDIA Blackwell GPUs to their limits, with code you can actually read and maintain. What we're showing at our booth State-of-the-Art GPU Performance with Mojo and MAX Live GPU programming on NVIDIA Blackwell: matmuls, generative AI model serving, and more. Porting a CUTLASS Blackwell Conv2D Kernel to Mojo See how Mojo's structured kernel architecture and AI-assisted development made it possible to port a complex CUDA C++ kernel in a single session, resulting in cleaner code that runs 6.6x faster than cuDNN on B200 GPUs. DeepSeek V3 on B200 via Modular Cloud DeepSeek V3 running live on NVIDIA B200 GPUs, served by the MAX stack with Mojo kernels. Real throughput and latency numbers for text and code generation workloads. FLUX.2-dev Image Generation on B200 Submit a prompt, get a high-quality image back fast. BFL's FLUX.2-dev diffusion model running on B200 with the full MAX serving stack, optimized end-to-end for throughput, latency, and cost per image. Why stop by? Stay in the loop RSVP to this event to receive updates on scheduled demo times and special sessions during GTC (and maybe a chance to meet Chris Lattner). We'll send you a heads-up before our most popular demos so you don't miss them. Find us on the expo floor Modular - Booth #3004 NVIDIA GTC | San Jose, CA Interested in a deeper conversation about deploying AI at scale? Book a meeting with our team.
Hoogtepunten
Writing high-performance GPU code is increasingly difficult due to rising TFLOPS benchmarks and evolving hardware. Modular aims to address this complexity, showcasing Mojo and MAX at NVIDIA GTC 2026 to demonstrate accessible and maintainable coding for NVIDIA Blackwell GPUs. The booth features state-of-the-art GPU performance, live programming, and tasks like matrix multiplications and generative AI model serving. The Mojo platform simplifies code development, such as porting a CUTLASS Blackwell Conv2D Kernel to a more efficient form with significant speed improvements. DeepSeek V3 and FLUX.2-dev illustrate real-world applications on B200 GPUs, showcasing text, code generation and image creation with optimized performance. Attendees are encouraged to RSVP for updates on demos and sessions, and find Modular at booth #3004.
Locatie evenement
Laat je netwerk weten dat je komt
Deel dit evenement om gesprekken te starten, collega`s uit te nodigen en van tevoren contact te leggen.
