Open-source phoneme-level English pronunciation assessment (Wav2Vec2 + DTW). Self-hosted alternative to Azure Pronunciation Assessment.
-
Updated
Aug 15, 2026 - Python
Open-source phoneme-level English pronunciation assessment (Wav2Vec2 + DTW). Self-hosted alternative to Azure Pronunciation Assessment.
Unity SDK for real-time Audio-to-3D facial animation powered by AI. Convert speech audio into expressive 3D facial blendshapes with a simple API.
canvas-based talking head model using viseme data
canvas-based talking head model using viseme data
Open-source example for integrating ElevenLabs conversational AI with animated avatars using Mascotbot SDK. Features real-time lip sync and natural voice interactions.
LAM Gaussian-Splat Avatars for AI voice agents. Realistic avatars for interview/support/sales agents.
A conversational 3D avatar engine built on Google's GNM Head model. Turns text into real-time, lip-synced, emotionally expressive 3D facial animation, rendered live in the browser with Three.js — no game engine required.
AI-powered lip sync sprite generator. Turn any face photo into 9 mouth shape sprites — and animate them in real time. By Editlingo (editlingo.com) · HistOracle (historacle.ai)
Real-time lip sync for 3D avatars, from any audio. Give it an avatar URL and an audio stream; a small model listens to the audio in the browser and drives the avatar's mouth. No text, no phoneme timings, no TTS-provider visemes, no server.
Playing Audio with lipsync using different Avathar expressions
Voice-LLM digital human orchestration: pluggable ASR/LLM/TTS providers, an energy VAD, a streaming pipeline, and a phoneme-to-viseme timeline generator with ARKit export. Zero runtime dependencies.
Spoken audio in, viseme timing and transparent frames out. Quickly and easily add cartoon style lips to your Lego, Barbie or animations. Exports to greenscreen, Final Cut Pro, GIF, etc. Use in Procreate Dreams, ToonSquid and other animation tools. Adds a version of Adobe Character Animator's lip-syncing capabilities.
FastAPI backend for a multilingual AI avatar system with text-to-speech and voice-to-voice translation. Integrates AWS Bedrock, Polly, Transcribe, and S3 for speech synthesis, transcription, and viseme mapping to enable real-time avatar lip-sync across multiple languages.
Turkish lip-sync (viseme) generator via forced alignment for TTS avatars — Türkçe TTS avatarları için dudak senkronu (viseme) üreteci
Reusable Azure Speech TTS and Viseme demo with Python SDK and HTTP clients, synchronized playback, raw evidence, and reproducible deployment.
Native Unity real-time lip-sync: transport-independent PCM playback plus HeadAudio-derived viseme analysis driving bilingual (English/Japanese) facial articulation. Unity 6.3, MIT.
Free, self-hosted 2D lip-synced avatars for LiveKit voice agents — no per-minute avatar fees.
To associate your repository with the viseme topic, visit your repo's landing page and select "manage topics."