ABSTRACT
01."Best" is measurable
Ask who has the best dubbing in the world and every vendor answers the same way. Search for the best dubbing software and the results all award themselves the crown. The word only means something once you decide what to measure. Dubbing has six dimensions, and each can be scored on the same clips:
- Spoken meaning preserved. Did the dub say what you said? Counted as important meaning mistakes per output, reviewed against one approved target meaning.
- Voice identity. Close your eyes: is this the same person? Measured as resemblance between the dubbed voice and the real speaker. YouTube auto-dubbing replaces your voice with a stock stranger voice; this dimension is where that shows up as a number.
- The face. In video, the mouth keeps moving in the source language unless the face is re-rendered: mouth, jaw, and expressions re-performed to match the new words. Audio-only tools skip this dimension entirely.
- Scene audio. Music, ambience, and sound effects should survive the dub; measured as spectral error between the dub's background and the source's.
- Vocal events and timing. Laughs, sighs, and beats should land where and how they landed; measured as event shape correlation and timing F1.
- Workflow and coverage. A tool that rejects your clip scores zero on that clip. Length minimums, batch caps, and language count are quality dimensions too.
One tiebreaker sits on top: latency. Two tools that tie on quality do not tie if one of them can dub a livestream while it airs; real-time is also the hardest setting, because nothing can be fixed in post. For how dubbing compares with subtitles and voice-over in the first place, see the video translation guide.
02.The method that makes comparisons honest
Paired and black-box: the same source clip and the same target language go to each system, and only the final user-facing audio is evaluated. Each finished dub is transcribed with word timestamps and reviewed against the same approved meaning, with valid paraphrases accepted. Acoustic measures compare the dub to the source scene. Every source clip carries equal weight, and 95% intervals resample source clips.
Two honesty rules matter most. First, coverage failures count as coverage, never as invented quality scores: a rejected clip is recorded as a rejection. Second, separate products are separate studies; results from different systems or versions are never pooled into one headline number.
03.The current numbers
We ran this method against ElevenLabs Dubbing twice on August 5, 2026: 418 paired outputs across 38 clips and 11 languages on the stable Dubbing v1 API, and 111 paired outputs across Mandarin, Spanish, and Japanese on their newest product, Website Dubbing V2 Alpha.
| Measure | Result | Raw scores |
|---|---|---|
| Review-flagged spoken-output mistakes | ElevenLabs had 95.5% more | 258 vs 132 |
| Background spectral error | ElevenLabs produced 261% more | 12.27 vs 3.40 dB |
| Laughter/reaction shape correlation | Familiar 243.7% higher | 0.759 vs 0.221 |
| Speaker resemblance | Familiar 27.8% higher | 0.461 vs 0.360 · all 11 languages favored Familiar |
| Sound-effect and beat timing | Familiar 24.3% higher F1 | 0.947 vs 0.762 |
| Predicted naturalness | Familiar 10.2% higher | exploratory |
| Measure | Result | Raw scores |
|---|---|---|
| Important spoken-meaning mistakes | ElevenLabs made 121% more | 64 vs 29 |
| Background-sound error | ElevenLabs produced 270% more | 11.92 vs 3.22 dB |
| Laughter/reaction loudness error | ElevenLabs produced 221% more | qualifying event regions |
| Laughter/reaction shape correlation | Familiar 28.2% higher | 0.960 vs 0.749 |
| Voice measures | Inconclusive | sample too small at 111 outputs; not pooled with the v1 study |
| Source-clip coverage | Familiar 38/38 · V2 Alpha 37/38 | V2 Alpha rejects clips under 11 seconds |
Familiar led on spoken meaning, scene audio, and vocal events in both studies, and on voice identity in every one of the 11 languages in the stable-API study; V2 Alpha's voice measures were inconclusive at this sample size, so that is how we report them. The one-line version we ship is the measured claim, not a slogan: "Translation quality and voice: tested more accurate than ElevenLabs." The clip-level breakdown is in the head-to-head paper; the full data and intervals are at the benchmark.
Translation text has its own study. On WMT24++ passages, Familiar's translations reached 96% of a first-pass professional translator's COMET score (95% CI: 95.2 to 96.9%), under a reference design that structurally favors the professional baseline. Per language: Hindi scored above the professional first pass, French tied it, and Indonesian reached 99.6%.
04.What nobody does well yet
Measured best is not solved. Fast-cut edits, memes layered over faces, and shaky footage are harder for our current Alpha, and some clips still deserve a human pass before publishing. No system in the field handles those scenes reliably today; they need a model that understands the whole scene, not just the face in it. Ours ships later this year. Until then, the honest answer to "is any dubbing perfect?" is no, and a benchmark that hides its hard cases is not a benchmark.
05.Run the test yourself
You do not have to trust anyone's numbers, including ours. The whole method compresses into an afternoon:
- Pick one clip of yourself with a laugh, background music, and a fast aside. Those three things break dubs.
- Dub it in two tools: same clip, same target language.
- Check the words. Have a native speaker, or a back-translation, compare the dub against what you actually said.
- Close your eyes and ask who is speaking. If the answer is "a narrator," voice identity failed.
- Listen for the room. Is the music still there? The ambience? Your laugh, and does it sound like yours?
- Watch the mouth. Does the face match the new words, or does the video contradict the audio?
Start by watching and listening to dubbed pairs at /demos. Testing costs little: Familiar's free tier dubs 4 minutes of video a month into 1 language with no card, and paid plans run $2.50 per finished minute per target language, all-in, across 25 languages, any to any (pricing). ElevenLabs quoted $3 at checkout on August 5, 2026, for audio only; published human dubbing packages run $39 to $94 per finished minute. For live streams, Familiar pushes each language to its own channel 10 to 30 seconds behind the source: the only voice + face translation in real-time. Run the clip, score the six dimensions, and the best AI dubbing tool for your channel stops being a marketing question.
REFERENCES
Dubbing is finally good. See the measurements, then try it on your own video.
