
AI Evals and Observability Strategies for Developers
Beschrijving
Building an AI agent for your application is only the beginning. The harder challenge is knowing whether it is actually working well once it starts planning, calling tools, making decisions, and operating across multiple steps. This workshop shows you how to design a practical evaluation and observability strategy for agentic systems. You’ll learn what to measure, how to evaluate agents during development, how to assess multi-step trajectories, and how to use production telemetry and online evals to detect failures and improve reliability. The goal is to help you move from “the agent seems to work” to a repeatable, evidence-based approach for measuring agent quality. NOTE: In case you're unable to purchase the ticket directly from Eventbrite, you may do so from Luma. Here's the link: AI Evals and Observability Strategies for Web Developers · Luma What you’ll learn By the end of the workshop, you’ll be able to: • Identify what should be evaluated in an agentic system • Design meaningful test cases and evaluation criteria • Run offline evals before deploying agent changes • Evaluate tool use, reasoning paths, and multi-step execution • Use traces and production signals to diagnose agent failures • Introduce online evals to continuously monitor quality • Build a repeatable evaluation strategy for your own agents What you’ll leave with You’ll receive: • A practical framework for evaluating AI agents • Guidance on choosing the right agent quality metrics • Approaches for offline evaluation during development • Techniques for assessing agent trajectories and intermediate decisions • A clearer understanding of observability for production agents • Strategies for combining offline and online evaluation • Best practices for building an eval strategy that evolves with your agent • Full workshop recording • Certificate of completion More importantly, you’ll leave with the vocabulary and intuition to discuss AI systems more precisely, investigate failures more effectively and make better decisions when building AI-powered products. Who should attend? This workshop is ideal for: • Web and Software Developers building agentic applications • AI Engineers • Platform and Backend Engineers • Technical Leads and Architects • Teams moving AI agents from prototype to production It is particularly useful for anyone who already has an agent or agentic workflow and wants a more systematic way to measure, debug, and improve it. Meet Your Instructor Supreet Kaur Sr. Gen AI Solutions Architect | AWS Supreet is a Senior GenAI Solutions Architect at AWS, helping startups take AI ideas from proof of concept to production. Her career spans data science, financial services, and cloud AI architecture, with roles at ZS Associates, Morgan Stanley, Microsoft, and AWS. She is the author of The AI Optimization Playbook, a speaker at 40+ events, a published thought leader, and co-inventor of a patented AI-powered personalization testing strategy.
Laat je netwerk weten dat je komt
Deel dit evenement om gesprekken te starten, collega`s uit te nodigen en van tevoren contact te leggen.