Додати подію
Hands-On Product Engineering with Claude Code
AIWorkshop

Hands-On Product Engineering with Claude Code

31 жов 202610:00 - 14:00ОнлайнEnglishOpen
LearnFind partners
TextDataCodeMusic

Опис

In this session you build an evaluation harness that defines correct before you write a line of agent code, redesign a single-context agent into an orchestrator with specialised subagents and tiered models, and measure the cost and quality difference between the two. You leave able to prove your agent reasons soundly rather than guessing well, and to point at the numbers that justify deployment. A workshop seat is a one-time investment. A poorly structured agent can burn 10x the value of that seat in unnecessary token usage at scale. Mike Hyzy is a Claude Certified Architect who builds agentic systems for top global companies, teaches AI strategy in Northwestern's MSAI program, runs innovation workshops featured at SXSW, and writes for Packt. What you will learn: During the workshop, you will learn how to: • Define what correct means before you build, using golden datasets, deterministic checks, a calibrated LLM judge, and trajectory scoring. • Evaluate agent behaviour as well as final answers, so you can distinguish sound reasoning from an answer that happens to be right by accident. • Build and diagnose a single-context Claude agent, using structured outputs and trajectory logs to identify where reliability breaks down. • Design an orchestrator-and-subagent architecture, with specialised roles, isolated context windows, restricted tool surfaces, and model tiering. • Measure the trade-off between quality and cost, and use evidence from the evaluation harness to decide which architectural changes genuinely improve the system. • Apply production controls, including permission rules, readable audit trails, telemetry, spend ceilings, and adversarial testing. • Make an evidence-based deployment decision, using measured accuracy, reasoning quality, failure modes, controls, cost, and residual risk. Who should attend this workshop: This is an advanced, hands-on masterclass for practitioners who have already built or shipped with Claude and are now facing the harder production questions: How do you know the agent is right? What is it allowed to do? What evidence would let someone approve it? It is best suited for: • product managers, product engineers responsible for agentic AI products or workflows that need to move beyond prototype. • Software and AI engineers building multi-step, tool-using, or multi-agent systems with Claude. • AI platform engineers standardising reusable agent patterns, evaluation methods, and controls across an organisation. • Engineering leaders who review or approve autonomous systems and need explicit, defensible criteria. • Founders and product teams whose product depends on agent reliability and who need more than a polished demo to justify deployment. This workshop is not suitable for beginners. It is not an introduction to prompting or basic agent building. Participants should already be comfortable using LLMs beyond a standard chat interface and following Python code. Build the Agent. Then Prove It. Building an agent is no longer the hard part. The harder problem is proving that a fluent, confident output is actually supported by a reliable process. Conventional software tests are not enough for analytical agents: several outputs may be valid, and a correct-looking answer can still be produced through invented reasoning. That leaves builders without evidence, approvers without criteria, and promising agents stuck in prototype. This masterclass flips the usual development sequence. You build the measurement apparatus first. Every architectural decision that follows is then tested against a fixed, visible standard. The practical system: a third-party vendor risk assessment agent. The agent ingests a vendor security and contract package - including a SOC 2 report, data processing agreement, master services agreement, and completed security questionnaire - and assesses it against an organisation’s control standard. It returns a structured risk memo with findings, severity, violated controls, evidence, and recommendations. The exercise is designed so that some findings can only be discovered by combining evidence across documents. That makes correct answers distinguishable from lucky answers and gives the evaluation harness something meaningful to measure. The same orchestration, evaluation, and control pattern transfers to other multi-document analytical workloads such as contract review, competitive intelligence, regulatory monitoring, and claims assessment. What you will leave with: You will leave with practical assets you can adapt to your own agentic AI projects: • A working repository you can point at your own document sets after the workshop. • A reusable method for defining correctness before development begins. • A scored evaluation harness combining a hand-written golden set, deterministic checks, an LLM judge, and reasoning-path evaluation. • An orchestrator-and-subagent architecture with isolated contexts and restricted tool access. • A practical production control layer covering permissions, auditability, cost limits, and failure handling. • A repeatable adversarial-testing approach for checking whether documents or inputs can manipulate the agent outside its intended boundaries. • A one-page deployment assessment carrying your own measured figures for accuracy, reasoning soundness, failure modes, controls, cost, and residual risk.

Інструменти, що будуть використовуватись на події

Claude

Розкажіть своїй мережі, що ви йдете

Поділіться цією подією, щоб розпочати розмови, запросити колег та налагодити контакти до її початку.

Теги

# AI# Workshop

Подібні події

Платформа штучного інтелекту

Ціни на сайті
Купити квитки