Native Unity real-time lip-sync component for Splatterface Games — a reference-exact C# port of HeadAudio (MFCC + Gaussian prototypes → 15 Oculus-style visemes) plus the streaming/timeline/facial layers needed to animate an authored 3D face in sync with the audio the player hears.
Built per HEADAUDIO_HANDOFF.md (+ JAPANESE_COVERAGE_ADDENDUM.md for
English/Japanese bilingual coverage) by eight parallel Devin CLI workers
(swe-2-max) in two waves; coordination contract in CONTRACTS.md. Results:
docs/acceptance-report.md.
| Path | What |
|---|---|
src/ |
SplatterfaceGames.LipSync.Core — netstandard2.1 analyzer port (no Unity deps), NUnit tests, CLI, vendored model |
unity-package/ |
com.splatterfacegames.lipsync — bounded PCM timeline, scheduled playback, viseme target queue, facial compositor, rig adapter, test face, diagnostics |
harness/ |
Node reference runner feeding fixtures through the upstream JS processor |
reference-output/ |
golden viseme-timeline/v1 JSON per clip |
fixtures/ |
12 deterministic SAPI speech WAVs + manifest (+extra/ real speech); fixtures/ja/ — 40 JSUT-based Japanese clips + manifest-ja.json |
ja/ |
Japanese analyzer routes: kana/mora parser, alignment route, HeadAudio-format acoustic model (model-ja-mixed.bin), route metrics |
vendor/ |
pinned clones: HeadAudio d3af5f9, uLipSync 0605879 (.git stripped; commits recorded in docs/provenance-*.md) |
testbed/ |
minimal Unity 6000.3.24f1 project: uLipSync baseline + package + core DLL + LipSyncTest.unity |
eval-output/ |
uLipSync timelines, comparison.md(-ja), editmode-results(-wave2).xml, ja/ ja-port/ ja-enbase/ route outputs |
docs/ |
per-worker reports + acceptance report |
workers/, logs/ |
worker prompts and session transcripts |
dotnet test src\SplatterfaceGames.LipSync.sln :: 19/19 — parity, chunk invariance
node harness\run-reference.mjs :: regenerate reference timelinesUnity side: see docs/eval-report.md → "Commands to reproduce" (batchmode
EditMode = 54/54 after wave 2, scene builder, uLipSync comparison) and
docs/ja-eval-report.md §7 for the Japanese validation entry points.
- Port parity: 100% label agreement, 0 ms boundary delta vs reference (5,381 English frames; 11,114/11,114 identical on the Japanese model too).
- HeadAudio > uLipSync for this use: uLipSync can't represent consonant closures and needs per-voice calibration; HeadAudio is pretrained with all 15 classes.
- Japanese: ship the alignment route for transcript-driven sessions — the
trained acoustic model works but the 6-frame vote erases brief onsets
(ja_FU 47.3%→7.1%); English model on Japanese audio = 33.2% agreement,
i.e. not coverage. Full matrix mapping in
docs/ja-eval-report.md. - Upstream limits carried into the port (documented): input must be ≥16 kHz
(precondition lower rates), sparse-event → dense-frame convention,
Weights/Confidenceare derived quantities, no tail flush. - Open gates: physical iPhone validation, human-marked contact ground truth,
real voice-transport integration, play-mode demo run,
en_Lproducer, authored mesh poses for the four reserved bilingual targets.
Code and original documentation: MIT (see LICENSE, copyright Jetha Chan).
Bundled third-party components carry their own licenses — see
THIRD_PARTY_NOTICES.md: HeadAudio + model-en-mixed.bin (MIT), uLipSync
(MIT), JSUT basic5000 audio + model-ja-mixed.bin (CC BY-SA 4.0, attribution
in fixtures/ja/NOTICE.md and ja/MODEL_LICENSE.md).