
Can Agents Actually Self-Improve?
Опис
An interactive evening constructing an agent that demonstrably improves: simulating edge cases, scoring against evaluations, and gating every release so the next version only ships if it surpasses the last. You'll leave with the loop, built on open-source tools, running on a laptop. Event details "Self-improving agents" is currently a hot topic in pitch decks. Some of it is genuine, much of it isn't, and this evening distinguishes the two by building the genuine version before you. We will collaboratively construct the authentic version live to showcase the differences. Before anything reaches production, we simulate synthetic users and adversarial scenarios to identify where the agent falters, score those attempts against evaluations, and reintegrate the failures. The evaluation becomes the gate: nothing is released unless the metrics move positively. Then we maintain it running, directing production traffic back through the same evaluations so regressions manifest as declining scores rather than support tickets. Full recursive self-improvement remains an open research issue and we won't pretend otherwise, but the constrained version is feasible today, on OpenTelemetry and open-source frameworks you can fork and self-host. Leading the workshop, Nikhil, founder and CEO of Future AGI, oversees the live build, from start to finish. He dedicates his days to the very problem this evening addresses: making agents dependable enough for production use, and measurable enough to confirm their actual improvement. Bring a laptop to: Instrument a baseline agent with OpenTelemetry and capture its initial evaluation scores Simulate synthetic users and adversarial scenarios to deliberately test its limits Score the attempts against evaluations to transform failures into metrics, not anecdotes Read the traces to discover the actual root cause, not to speculate from the output Reinstate the solutions, then re-simulate to verify the scores have improved Gate the release: promote a new version only if it exceeds the previous one Route production traffic back through the same evaluations to ensure continuous improvement Agenda 5:30 - Snacks & Greetings 6:00 - Guest speaker (to be confirmed) 7:00 - Live build: the self-improvement loop that truly ships (bring a laptop) 7:30 - Open debate & Q&A: hype versus reality 8:00 - Networking 8:30 - See you in the next iteration For engineers & AI product developers managing agents in production who seek the version that indisputably improves, built on open tools they can maintain. About Pebblebed Pebblebed is a technical early-stage VC established by Pam Vagata (co-founder of OpenAI, previously ran AI for Stripe, inventor of FBLearner Flow); Keith Adams (founder of Facebook AI Research, former chief architect at Slack, 20th engineer at VMWare) & Tammie Siew (former Sequoia Southeast Asia investor, previous Sequoia & Notable Capital backed founder) About Future AGI Future AGI is an open-source AI simulation, evaluation, and observability platform. Teams utilise it to simulate agents before they are deployed, assess them against real failure modes, and continue monitoring them in production to ensure quality doesn't gradually diminish. It's self-hostable and OpenTelemetry-native, with tracing that integrates into 35+ frameworks. The loop you'll construct tonight operates on the same open tools, yours to fork and take home, no account necessary.
Місце проведення
Розкажіть своїй мережі, що ви йдете
Поділіться цією подією, щоб розпочати розмови, запросити колег та налагодити контакти до її початку.







