
NYC Apache Iceberg™ Community Meetup
Description
🧊 Apache Iceberg returns to the Big Apple! 🗽 We’re bringing the Apache Iceberg™ community back to New York for a day of technical talks, practical insights, and conversations with the people shaping the future of open data infrastructure. Whether you’re running Iceberg in production, evaluating it for your lakehouse, or exploring the ecosystem, this meetup is a chance to learn from practitioners, connect with fellow data engineers, and hear real-world lessons from modern data teams. Date: August 20th Interested in speaking? 🎤 Submit your CFP here: https://forms.gle/K3SsWXpXJKgkrPc88 Location: Amazon JFK27 - Hank Agenda• 2:00 PM – 2:30 PM: Doors Open, Snacks & Networking Talk Descriptions:Ali Alemi, AWS — Pr. Specialist Solutions ArchitectSub-Second Fraud Detection Meets Agentic AI: Streaming, Iceberg, and Automated ForensicsA payment is flagged as fraud 800 milliseconds after it hits an account, and within five seconds an AI agent reverts it and autonomously investigates six months of history, tracing the actor's footprint and delivering a compliance-ready forensic report — with no human involved. The talk follows that ten-second journey to show how real-time data and Apache Iceberg give agentic AI the speed and depth to not just detect fraud but resolve it, from instant remediation to automated investigation. It closes on the hardest design question: when the agent should act on its own, and when it should call in a human. Jayce Slesar, BETA Technologies — Data Infrastructure EngineerIceberg for Massive Aircraft TelemetryAn overview of how BETA Technologies uses Iceberg and Trino to query trillions of data points at lightning-fast speeds, meeting the needs of dozens of teams across the organization who analyze flight data. Conor McCarter, Prequel — Co-FounderLet's Get Physical: Accelerating Iceberg Table ReplicationOpen table formats like Iceberg give independent engines a shared rulebook, which unlocks a third replication option beyond proprietary physical replication and engine-based logical replication: non-proprietary, physical replication. Naively this looks easy — pick a snapshot, enumerate its files from the manifests, copy them in parallel, and publish a new table — but operating beneath the query engine raises hard questions about authorization in multi-tenant lakes and whether you've truly reproduced the source table's logical state. The talk walks through Prequel's cross-cloud/cross-region implementation and the Iceberg features that complicate it (partition transforms and evolution, positional vs. equality deletes, schema evolution leaving dropped columns in Parquet, historical metadata exposure), sharing the architecture that emerged and leaving attendees with a practical checklist for building directly against Iceberg storage. About PuppyGraph PuppyGraph is the first and only real time, zero-ETL graph query engine in the market, empowering companies to transform existing relational data stores into a unified graph model in under 10 minutes, bypassing traditional graph databases' cost, latency, and maintenance hurdles. 💬 Join PuppyGraph Community Slack 📚 Check out PuppyGraph Engineering Blog 📲 Follow PuppyGraph on LinkedIn & Twitter 🖥️ Subscribe to PuppyGraph YouTube 💾 Download PuppyGraph Forever Free Developer Edition (no form & no payment required) About AWS Whether you're looking for generative AI, compute power, database storage, content delivery, or other functionality, AWS has the services to help you build sophisticated applications with increased flexibility, scalability, and reliability. AWS is the world's largest Cloud Services provider. https://aws.amazon.com/ At AWS, Apache Iceberg is an open-source table format that simplifies table management while improving performance. AWS analytics services such as Amazon SageMaker Lakehouse, Amazon S3Tables, Amazon EMR, Amazon Glue, Amazon Athena, and Amazon Redshift include native support for Apache Iceberg, so you can easily build transactional data lakes on top of Amazon Simple Storage Service (Amazon S3) on AWS. Additional Resources and Information: 📚 Workshop: Running Apache Iceberg on AWS 📚 Blogs: Apache Iceberg on AWS 📚 AWS Prescriptive Guidance: Using Apache Iceberg on AWS 🖥️ Subscribe to AWS Events and AWS Developers 💜 We’re hiring, join our team About Microsoft Fabric As AI reshapes every industry, one truth remains constant: data is no longer just an asset—it’s your competitive edge. The pace of AI demands easy data access, faster insights, and the ability to iterate without friction. Yet many organizations are held back by fragmented data estates and legacy systems. Microsoft Fabric was designed to meet this moment—to unify your data, simplify your architecture, and accelerate your path to becoming an AI-led organization. Apache Iceberg is the connective tissue enabling native and 3rd parties to integrate and interoperate in a seamless manner without forcing users to build or maintain complex lake management solutions. Fabric OneLake is a single place to store and access data in Iceberg and Delta Lake formats. Additionally, it provides a unified catalog powered by Iceberg REST specification enabling seamless integration with your favorite tools like Fabric native engines, Databricks, Snowflake, ClickHouse and many more. Introducing Microsoft Fabric Introducing Microsoft OneLake Learn more about Iceberg in Microsoft Fabric Connect to OneLake using Iceberg REST-powered Table APIs About ClickHouse Established in 2009, ClickHouse leads the industry with its open-source column-oriented database system, driven by the vision of becoming the fastest OLAP database globally. The company empowers users to generate real-time analytical reports through SQL queries, emphasizing speed in managing escalating data volumes. 📚 Get started on ClickHouse Github 💬 Join the ClickHouse Slack Channel 💜 We’re hiring, join our team!
Event location
Let your network know you`re going
Share this event to start conversations, invite colleagues, and connect before it begins.