Voice Agent in Foundry Agent Service offers below key values to customers:
Build and launch an enterprise-ready voice agent in under two minutes. Choose speech-to-speech or cascaded pipelines powered by OpenAI, Microsoft AI, and Azure real-time models. Extend your agent with Foundry tools, monitor and measure its performance, and connect inbound and outbound calls through Teams and Twilio. Create engaging conversations with voices optimized for call centers and lifelike avatars.
This repository is the official community hub for Voice Agent in Foundry Agent Service. Here you'll find:
🐛 Report Issues — File bugs, feature requests, and feedback via GitHub Issues
📚 Resources — Curated links to docs, videos, blogs, and community content for Voice Agent
🧪 Samples — Hands-on samples and extended solutions
Voice Live is the foundation for real-time voice interaction, and Voice Agent builds on that foundation to deliver complete voice-first agents. Voice Live handles the real-time voice experience, while Voice Agent adds the intelligence, tools, knowledge, and orchestration required to build end-to-end agentic applications.
For most customers building voice-first agents, we recommend starting with Voice Agent. Voice Agent provides the more complete, integrated experience for building and operating an agent, bringing together voice, reasoning, instructions, knowledge, tools, and orchestration. Voice Live is a better fit when customers already have their own agent stack and primarily need real-time voice capabilities with greater control over the voice application architecture.
| Resource | Description |
|---|---|
| Product home page | Explore Foundry Agent Service, including voice capabilities. |
| Foundry portal | Open your Foundry project. |
| New Foundry portal — Create & manage agents | Create and manage voice agents in the new Foundry portal. |
| Documentation & quickstart | Create a voice-based prompt agent, customize its instructions, and test spoken conversations. |
| Configure a voice agent | Configure your voice agent's behavior and voice settings. |
| Use a hosted agent as the conversation engine | Use a hosted agent for conversation logic and tools, while Voice Live handles speech, turn-taking, and interruptions. |
| Use a subagent in a voice-based agent | Delegate specialized requests to prompt or hosted subagents in the same Foundry project. |
| Voice Agent tracing, monitoring & evaluation | Monitor voice sessions and evaluate conversation transcripts using datasets, traces, or simulations. |
| Pricing & billing | Understand Voice Agent pricing and billing. |
| Foundry blogs | Browse the shared Microsoft Foundry blog for announcements and tutorials. |
| Voice Live video tutorial | Related Microsoft Learn training on building a Voice Live agent in Foundry. |
| GitHub — Official resources & samples | Find Voice Agent resources and runnable examples. |
| Report issues & feature requests | Report bugs, request features, and share feedback. |
| Hands-on samples | Explore runnable samples and extended solutions. |
| Tools & integrations — MCP examples | Explore the shared MCP service and tool integration examples. |
| Regions & quotas | Check the shared Foundry Agent Service region and quota documentation. |
| Tech Community discussions | Join the shared Microsoft Foundry community forum. |
| Contact the Voice Agent team | Email [email protected] with questions and feedback. |
| Python SDK | azure-ai-projects for agent management and realtime voice sessions. |
| Java SDK | com.azure:azure-ai-agents for voice sessions, conversations, and telephony. |
| JavaScript / TypeScript SDK | @azure/ai-projects for agent management and realtime voice sessions. |
| C# / .NET SDK | Azure.AI.Projects.Agents for voice agent definitions and management. |
In the Foundry portal, open your project, go to Build > Agents, select Build an agent, and choose Voice as the interaction mode.
Voice Agent is now available in public preview. Customers can start building voice agents and preparing for production.
Explore Voice Agent in Foundry Agent Service for voice agent guides, portal links, and runnable samples.
- 2026.09 Ship agents faster with expanded model choice, voice agents, and continuous optimization — Overview of the latest Foundry Agents capabilities, including a Voice Agent video overview.
- 2026.09 Introducing voice agents in Microsoft Foundry — Voice Agent public preview announcement.
- 2026.06 Azure Speech at Build 2026: Powering Voice Agents with Real-Time and Life-like Experiences — Azure Speech updates from Build 2026 for real-time voice agents and lifelike experiences.
- 2025.11 Advancing Speech Innovation with Azure Speech in Microsoft Foundry — Azure Speech innovations in Microsoft Foundry for conversational AI and voice agents.
- 2025.09 Upgrade your voice agent with Azure AI Voice Live API — General availability announcement for Voice Live API, enabling real-time speech-to-speech voice agents.
- Astra Tech brings Voice Live API in Azure AI Foundry to its fintech-first app — Astra Tech uses Voice Live in botim for a multilingual voice assistant that helps users complete tasks such as international money transfers. The story reports 300,000 monthly active users and 100,000 daily active users.
- Boosting patient satisfaction with healow Genie and Voice Live API in Azure AI Foundry — The article describes a Voice Live pilot for healow Genie, covering appointment information, common questions, and voicemail callbacks. Pilot: It discusses anticipated benefits rather than measured improvements from a completed rollout.
- Kansai Television: AI Hachiemon — Microsoft AI Co-Innovation case study. Kansai Television built a conversational version of its mascot, Hachiemon, using Voice Live and Azure AI Foundry. Azure AI Speech Custom Voice recreates the character's distinctive voice for entertainment and interactive experiences.
Speech-to-speech models can achieve latency as low as 500 ms on service side. Cascaded pipelines (STT → LLM → TTS) can also deliver low latency < 1s with the right LLM and streaming speech configuration. Actual latency depends on the model, region, network conditions, and tool calls.
Voice Agent brings together advanced speech and language models, including GPT Realtime, GPT Live, Azure speech-to-text and HD voices, and MAI Transcribe and MAI Voice. Choose the combination that best fits your languages, domain, and conversational experience.
Recent speech benchmark references:
- 2026.10 MAI-Transcribe-2-Streaming benchmark results — Microsoft reports first place for final and partial transcript accuracy on Artificial Analysis in its October 1 announcement.
- 2026.09 MAI-Transcribe-2 benchmark results — Microsoft reports first place on FLEURS across 60 languages and second place on the Artificial Analysis word error rate leaderboard in its September 3 announcement.
- 2026.06 Azure LLM Speech benchmark results at Build 2026 — Microsoft reports first place on the Open ASR Leaderboard for the updated LLM Speech model in its Build 2026 announcement.
- Speech-to-speech — Azure Realtime: Custom voices are available upon request. Contact [email protected] to discuss access. See the Azure Realtime voice configuration reference for the publicly documented native voice settings.
- Cascaded pipelines — speech recognition and synthesis: Adapt recognition to your vocabulary and audio conditions with Azure Custom Speech. Create a distinctive voice with Azure Custom Voice, or use MAI Voice customization to create a voice from a short reference recording through the documented access and consent workflow.
Voice Agent can scale to meet your business needs. Customers are already using it to run more than 3,000 concurrent sessions and are continuing to scale. For large-scale deployments with high concurrency requirements, contact us to discuss your capacity and scalability needs.
This directory is the entry point for the Voice Agent portal, runnable samples, the Finance reference workflow, and its shared MCP service. Detailed setup and operation instructions live with the component that owns them; this file routes users and coding agents to the correct entry point.
Important
On Windows, use WSL2 for this repository. Do not run the repository
setup or runtime workflow from native PowerShell or Command Prompt, and do
not use a checkout mounted under /mnt/c/. Start WSL2, clone the repository
into the WSL Linux filesystem (for example, ~/src/voice-agent),
and run all repository commands from WSL. Use the Windows browser to open
the resulting localhost UI and grant microphone permission.
The standalone audio rewrite and translation sample also supports native Windows PowerShell with Python 3.11+, without WSL.
For an unqualified request such as "run the UI" or "start the portal",
use portal/. It is the general Voice Agent UI and runs on
http://127.0.0.1:9527 by default. Do not start the Finance MCP stack unless
the request mentions Finance, Templates, or the shared MCP.
Use this path when the portal must publish or run the checked-in Templates. It is the recommended end-to-end Linux/WSL workflow:
cd /path/to/voice-agent
az login
./scripts/setup-local-examples.sh \
--project-endpoint "https://<account>.services.ai.azure.com/api/projects/<project>"
./scripts/login-devtunnel.sh
./scripts/setup-local-examples.sh --check
./scripts/manage-local-mcp-and-ui.sh restartlogin-devtunnel.sh uses GitHub device-code authentication. Follow the printed
https://github.com/login/device prompt. Azure operations continue to use the
separate identity selected by az login.
Success requires local_mcp_and_portal=ready. Open
http://localhost:18098, then verify with:
./scripts/manage-local-mcp-and-ui.sh status
curl -fsS http://127.0.0.1:18003/healthz
curl -fsS http://127.0.0.1:18098/healthzPort 9527 is only for a portal-only session without the managed local MCP
lifecycle. Do not run a second manual portal after this quickstart.
| Goal | Start here | What it owns |
|---|---|---|
| Run the general Voice Agent UI | portal/README.md |
Agent editor, YAML version editing, Templates, voice playground, and standalone WebRTC page |
| Run Python or .NET samples | samples/README.md |
Common Python setup, microphone samples, REST lifecycle, IQ, Toolbox, local functions, downloads, and the C# sample |
| Run the GPT Live terminal sample | samples/gpt_live/README.md |
Create or reuse a GPT Live agent, stream microphone audio with the OpenAI SDK, and view independently scrollable GPT Live, delegation, and user transcripts |
| Run the complete Finance workflow | docs/README.md |
Ordered subscription, MCP, sample, portal, and debugging guides |
| Work on or deploy the Finance MCP | shared_mcp/README.md |
Shared MCP image, Finance routes, local Dev Tunnel hosting, and Azure Container Apps deployment |
| Create an IQ + voice + avatar Agent | samples/create-agent-with-iq-avatar-voice/README.md |
Portal-first Andrew Dragon HD, Harry Business, Knowledge IQ, and optional Python creation |
| Inspect SDK package information | dist/README.md |
Public Python and .NET SDK dependencies and historical preview build records |
| Use the coding-agent workflows | skills/ |
Voice Agent creation, IQ/Toolbox provisioning, and local-session debugging |
When the user asks to run or debug something from this directory:
- Verify that commands will run on Linux. For a Windows user, require a WSL2
checkout in the WSL filesystem. If the checkout is under
/mnt/c/or the terminal is native Windows, stop and guide the user to clone and reopen the repository in WSL2 before continuing. Exception: the standalonesamples/audio_file_rewrite/sample supports native Windows; follow its README. - Select the component from the table above and read its
README.mdbefore running commands. - Treat UI without a qualifier as the general
portal/. - Treat Finance UI, portal Templates with Finance, or
shared MCP UI as the workflow documented in
docs/03_run_samples.md. - Reuse an existing component-local
.envand virtual environment when they are valid. Never copy credentials or endpoints between unrelated.envfiles without the user's intent. - Start servers as long-running processes, verify their
/healthzendpoint, and report the browser URL and log location. Do not report success merely because a process was spawned. - Do not silently fall back to mock data or a different Azure Project when authentication, endpoint, model, or preview checks fail.
For the default portal route, follow portal/README.md to
prepare portal/.env and its .venv, start portal/demo_server.py, then
verify:
GET http://127.0.0.1:9527/healthz
For the complete Finance route, follow the ordered
docs/README.md workflow. The normal lifecycle commands are:
./scripts/setup-local-examples.sh --project-endpoint "https://<account>.services.ai.azure.com/api/projects/<project>"
./scripts/manage-local-mcp-and-ui.sh restart
./scripts/manage-local-mcp-and-ui.sh status
./scripts/manage-local-mcp-and-ui.sh stopThe setup command installs Node.js 22 and the Dev Tunnel CLI under the ignored
.local-mcp-and-ui/tools/ directory when they are missing. It also supports
minimal Linux Python installations that provide venv but omit ensurepip:
each component environment receives its own bootstrapped pip. No sudo or
system package change is required for those three tools. Azure CLI must already
be installed and authenticated; Dev Tunnel device-code authentication remains
an explicit user action because setup must not authenticate as an identity the
user did not choose.
The supported repository working environment is Linux. On a Windows computer,
use WSL2 for the entire repository workflow, including the portal, samples,
MCP, deployment, and validation commands. The standalone
audio_file_rewrite sample is an exception
and can run and be tested directly in Windows PowerShell:
- Start a supported WSL2 Linux distribution.
- Clone this repository again into the WSL filesystem, for example under
~/src/. Do not run the workflow from a Windows checkout mounted under/mnt/c/. - Install Git, Python 3.10+, and Azure CLI inside WSL, then authenticate Azure CLI. The local setup script can install repository-local Node.js 22 and Dev Tunnel CLI copies. Azure Developer CLI is needed only for deployment. Windows-side CLI login state is not assumed to be shared.
- The local setup, MCP E2E, and lifecycle scripts use native Python and never
invoke Docker. Install Docker only when deliberately running the separate
shared_mcp/scripts/package.shimage-packaging command; Azure Container Apps deployment uses a remote build. - Open the WSL checkout with VS Code Remote - WSL and run all commands from its WSL terminal.
- Open the resulting
localhostURL in the Windows browser; browser microphone permission remains on the Windows side.
Some individual component documents retain native PowerShell commands because
their code can run independently on Windows. They are not the recommended or
supported end-to-end repository workflow. Except for the standalone sample above,
native Windows execution, WSL1, and
running the checkout from /mnt/c/ are outside the supported path.
| Path | Purpose |
|---|---|
portal/ |
General local Voice Agent portal and WebRTC UI |
samples/ |
Python, .NET, and Finance samples |
docs/ |
Finance architecture, setup, operation, and debugging guides |
shared_mcp/ |
Shared Finance MCP source, native runtime, optional container, IaC, and deployment scripts |
scripts/ |
Finance local setup and process lifecycle entry points |
dist/ |
Public SDK package information and historical preview build records |
skills/ |
Reusable coding-agent workflows |
tests/ |
Offline contracts for the common Projects SDK samples |
- Configure only Azure resources and identities the customer is authorized to use.
- Never place access tokens, API keys, connection secrets, or customer data in source files.
- Use Foundry Project connections or environment-based credentials.
- The portal and samples use Azure Foundry; running the local portal does not run the Azure voice service locally.
- Review each component's persistence and recording behavior before using sensitive prompts, audio, transcripts, or tool output.