HEAD-TO-HEAD BENCHMARK

Familiar vs YouTube Auto-Dubbing: A Frozen-Cohort Study

FAMILIAR RESEARCHALL ARTICLES

ABSTRACT

We compared complete delivered dubs from Familiar and YouTube's auto-dubbing on the same source clips: 50 selected clip-language pairs across Arabic, Spanish, French, Hindi, and Korean, frozen on August 8, 2026. Familiar led on speaker resemblance (+182.9%), important meaning kept (70.3% fewer ideas changed or lost), background sound, and vocal events. A usable auto-dub existed for 25 of 43 source clips; Familiar dubbed all 43. Full data and six synchronized listening examples: thefamiliarlab.com/benchmark/youtube.

01.Method

  • Paired, black-box: one source clip × one target language; complete final audio from both systems.
  • Timing-sensitive metrics: only the 31 pairs with independently verified source-timeline alignment (40 ms contract; uncertain lanes fail closed).
  • Event metrics: the 14 lanes with human-verified laughter or reaction regions.
  • Unavailable YouTube lanes: stay explicit, never scored as zero.
  • Meaning review: judges the delivered speech; valid paraphrases and tiny background reactions are not errors; CER/WER excluded.

02.The voice: theirs is a stranger's

YouTube's auto-dub speaks in a generic synthetic voice; Familiar re-renders yours. On the 31 exactly aligned pairs, speaker resemblance to the real speaker:

FIG. 01 · SPEAKER RESEMBLANCE (SIM-O) · 31 ALIGNED DUBS · FAMILIAR +182.9%
FAMILIAR0.399
YOUTUBE0.141

03.The meaning

Of 528 important source ideas across the 50 reviewed dubs, YouTube changed or lost 145; Familiar 43.

FIG. 02 · IMPORTANT SOURCE IDEAS CHANGED OR LOST · OF 528 · FAMILIAR −70.3%
YOUTUBE145
FAMILIAR43
34/50

YouTube dubs with an important spoken-meaning issue

18/50

Familiar dubs with an important spoken-meaning issue

12 vs 0

critical meaning failures, YouTube vs Familiar

04.The scene and the reactions

FIG. 03 · BACKGROUND SPECTRAL ERROR · SPEECH-FREE REGIONS · YOUTUBE +232.6%
YOUTUBE11.21 dB
FAMILIAR3.37 dB
0.998 vs 0.236

laughter/reaction shape correlation · 14 event lanes

0.11 vs 3.36 dB

laughter/reaction loudness error · 14 event lanes

25/43

source clips with a usable YouTube auto-dub — Familiar dubbed all 43

05.Limits

REFERENCES

  1. [1]Familiar vs YouTube Auto-Dub: full results, method, and listening examples
  2. [2]All Familiar benchmarks: ElevenLabs, YouTube, human translators
  3. [3]YouTube Help: Use auto dubbing
  4. [4]YouTube Auto-Dubbing Review: What We Measured

Dubbing is finally good. See the measurements, then try it on your own video.