[TVCG] Character-specific Fine-tuning (CsF) and temporal inference from AnyTalk
-
Updated
Aug 16, 2026 - Python
[TVCG] Character-specific Fine-tuning (CsF) and temporal inference from AnyTalk
Real-time lip sync for 3D avatars, from any audio. Give it an avatar URL and an audio stream; a small model listens to the audio in the browser and drives the avatar's mouth. No text, no phoneme timings, no TTS-provider visemes, no server.
Native Unity real-time lip-sync: transport-independent PCM playback plus HeadAudio-derived viseme analysis driving bilingual (English/Japanese) facial articulation. Unity 6.3, MIT.
🎙️ Apple-grade Voice User Interface (VUI) animation generator with sample-accurate audio sync and transparent broadcast export (ProRes 4444, WebM Alpha, PNG Sequence, Chroma MP4).
Real-time, muscle-aware multilingual lip sync for robots and constrained displays — English, Mandarin, and Spanish
An AI-powered pipeline that transforms text into realistic lip-synced talking face videos using ElevenLabs and Wav2Lip.
To associate your repository with the speech-animation topic, visit your repo's landing page and select "manage topics."