ABSTRACT
01.What AI lip sync is
A lip-sync model takes finished audio in the new language and warps the mouth region to match it. That solves the most visible mismatch of classic dubbing, lips speaking the old language, and nothing else.
- Lips: re-timed to the new audio.
- Voice: not lip sync's job; where most dubs lose the person.
- The rest of the face: still performing the old language, half a beat off.
- The scene: music, effects, laughter; untouched or damaged upstream.
02.Why lips alone fall short, measured
The failure people actually notice is that the dub stops being the person. All numbers below come from paired studies on the same clips, published in full at /benchmark.
more important meaning mistakes from ElevenLabs (64 vs 29)
THEIR NEWEST DUBBING V2 (ALPHA) · 111 PAIRED DUBS · 3 LANGUAGESmore background-sound damage from ElevenLabs (11.92 vs 3.22 dB)
SAME STUDY · MUSIC, EFFECTS, AND ROOM AMBIENCEof expert human translation quality from our production pipeline
FRENCH, CHINESE, HINDI, INDONESIAN (FIRST PASS FOR EXPERT PROFESSIONAL TRANSLATORS)A perfect lip sync on a stranger's voice with a damaged soundtrack is still a stranger in your video.
03.Lip sync and everything past it
Familiar generates the voice and the face together, so the lip sync comes built in and the edit extends past the lips: eyes, brow, cheeks, and head re-perform the sentence in the new language, in your voice, with the music and the laughs where they were. Recorded video runs $2.50 per finished minute (ElevenLabs quoted $3 at checkout, audio only, August 5, 2026); livestreams run in real time at 42¢ per minute per language, the only voice + face translation that works live.
REFERENCES
Dubbing is finally good. See the measurements, then try it on your own video.
