Submit Event
Structure of Voice AI Agent

Structure of Voice AI Agent

05 Aug 202618:30 - 20:30 Asia/RiyadhRiyadh, Saudi Arabia10 AttendeesOpen

Description

About the workshop What actually happens from the moment you start talking to a smart voice agent until you hear its response? This session takes us on a technical tour inside the architecture that underpins real-time voice AI agents, to understand how their components work together as a unified system: from speech recognition and conversion to text (STT), to processing requests using large language models (LLMs), and then generating the voice response via text-to-speech (TTS) technologies. The experience does not stop at these three components; the session also addresses one of the most complex aspects of voice systems: how the agent knows when to listen, when to speak, how to handle interruptions and turn-taking, and how to reduce response latency to achieve a conversation that feels natural and smooth instead of a series of separate requests and responses. Then we move on to the toughest challenge: building these systems for the Arab world. Arabic does not mean a model that handles one language in a fixed form; rather, it is a linguistic system where dialects intertwine, pronunciation patterns and vocabulary differ, and code-switching occurs between Arabic and English within the same conversation. These characteristics impose additional challenges at various stages of the system, from accurate speech recognition to understanding context and generating natural-sounding voice that suits the user. These challenges increase when transferring the voice agent from a prototype model to an actual product operating within regulated Gulf sectors, such as banking and healthcare; where the quality of the model alone is not enough, but considerations of privacy, reliability, compliance, and handling user data must be embedded within the core design of the system. The session dissects these layers to understand what is required to build an Arabic voice agent capable of functioning in a real environment, and not just providing an impressive technical experience. About the speaker Fares began his career as an AI engineer, before transitioning into product and business roles, combining his technical background with an understanding of building and operating AI products in the market. He currently leads product operations at Surge, an AI platform that is primarily built for Arabic (Arabic-first), working on developing voice and text AI agents for regulated sectors in the Gulf region. Through his work at the intersection of AI, products, and operations, Fares focuses on transforming AI technologies from technical models into usable and operable products within real work environments and their regulatory requirements.

Event location

Let your network know you're going

Share this event to start conversations, invite colleagues, and connect before it begins.

about_us.editorial_intelligence_platform

Free