ABSTRACT
01.What AI video dubbing is
Dubbing replaces the spoken track of a video with a performed translation. AI video dubbing generates that performance instead of hiring an actor per language: the video stays exactly as shot, while the speech is translated and re-performed. Done right, the voice stays yours, and the face is re-rendered so the mouth, jaw, and expressions match the new words.
That last part separates dubbing from the alternatives: subtitles ask the audience to read, and voice-over hands your delivery to a narrator (the trade-offs are compared in Video Translation: The Complete Guide). The search keyword for the voice half is often "voice cloning", but identity is the real requirement: someone who knows you should watch the dub and recognize the same person. Finished examples, original and dub side by side, are at /demos.
02.How it works, step by step
A modern pipeline runs five stages on every clip:
- Speech recognition. The source audio is transcribed with word timestamps, so the system knows what was said and exactly when.
- Translation. Good dubbing translation rewrites idioms and jokes so they land in the target language, never word for word. A Do Not Translate list keeps catchphrases and names exactly as you say them.
- Voice generation. The translated line is generated as your voice, with your delivery, not as a stock narrator.
- Face re-performance. The face is re-rendered so mouth, jaw, and expressions match the new language, not left running under audio it never spoke.
- Scene audio. Music, sound effects, and room ambience are preserved underneath the new speech.
Between translation and render, Familiar pauses each video at a review screen: any translated line can be edited before the render starts.
03.The stranger-voice problem
Most dubs fail the identity test, and the failure has a structure. When the voice comes from one tool and the face, if it is touched at all, from another, nothing binds the output to the person: the generated voice drifts toward a generic register, the unchanged face contradicts the new audio, and the scene gets stripped because the pipeline isolates speech by discarding everything around it. YouTube auto-dubbing replaces your voice with a stock stranger voice. Anyone who has clicked a translated audio track knows the result — the picture is you, the voice is not.
This is measurable. In a same-clips audit (August 5, 2026), ElevenLabs Dubbing dropped or invented words 45 times, broke laughs and sound effects 8 times, lost the speaker's voice 6 times, collided overlapping speakers 6 times, flattened the delivery 4 times, stripped the scene ambience 4 times, and refused clips under 11 seconds outright.
Paired studies on the same source clips show the same pattern. Against the stable ElevenLabs Dubbing v1 API (418 paired outputs, 11 languages), Familiar measured 27.8% higher speaker resemblance, with all 11 languages favoring Familiar, and made 48.8% fewer review-flagged spoken-output mistakes (132 vs 258). Against the newest product, Website Dubbing V2 Alpha (111 paired outputs, Mandarin, Spanish, and Japanese), ElevenLabs made 121% more important spoken-meaning mistakes (64 vs 29) and produced 270% more background-sound error (11.92 vs 3.22 dB). The full head-to-head is in Familiar vs ElevenLabs Dubbing, with data and listening examples at /benchmark.
04.What to check before trusting a tool with your channel
Run one short clip through any tool first, and check six things:
- The eyes-closed test. Play the dub without the picture. Someone who knows the speaker should still say it's them.
- The face. Mouth, jaw, and expressions should move with the new words, not stay frozen under audio the face never spoke.
- The background. Music, sound effects, and room ambience should survive under the new speech, not vanish with the old one.
- Length and batch limits. Some products refuse short clips outright; ElevenLabs rejects anything under 11 seconds (observed August 5, 2026).
- The real price of a finished minute. Ask what one delivered minute costs with translation, voice, face, and scene audio all included, not a base rate with add-ons.
- Correction before publish. Can you edit a translated line before it renders? Familiar pauses every video at a review screen for exactly this.
05.What it costs
| Option | Per minute | What it covers |
|---|---|---|
| Familiar | $2.50 | All-in: translation, your voice, scene audio preserved, the face re-rendered |
| ElevenLabs Dubbing | $3 (checkout quote, Aug 5, 2026) | Audio only |
| Human dubbing packages | $39 to $94 | Recorded voice actors, audio only |
Familiar covers 25 languages, any to any. The free tier dubs 4 minutes of video a month into 1 language, no card; paid plans start at $35/mo and include all 24 target languages. Full tiers are on /pricing.
06.Where this is going
Dubbing a finished upload is the asynchronous case. The same pipeline now runs live: a stream is dubbed as it airs, and each target language is pushed to its own channel 10 to 30 seconds behind the source. Familiar is the only voice + face translation in real-time; how live dubbing works, and what it costs, is in Real-Time Dubbing for Livestreams.
Our current model is Familiar Alpha; a model built to understand the full scene ships later this year.
REFERENCES
Dubbing is finally good. See the measurements, then try it on your own video.
