ABSTRACT
01.Method
Both studies are paired and black-box: the same English source clip and target language went to each system, and only the final user-facing audio was evaluated. Each finished dub was transcribed with word timestamps and manually reviewed against the same approved target meaning; valid paraphrases were accepted. Acoustic measures compare the dub to the source scene. Every source clip carries equal weight, and 95% intervals resample source clips. The two cohorts are reported side by side and never pooled.
Coverage failures count as coverage, not as fabricated quality scores: when V2 Alpha rejected the 10.94-second clip, it was recorded as 37/38 coverage, and no synthetic audio-quality value was assigned.
02.Results: ElevenLabs Website Dubbing V2 Alpha (their newest product)
111 paired final-audio outputs across Mandarin, Spanish, and Japanese, on the 37 clips ElevenLabs accepted.
| Measure | Result | Raw scores |
|---|---|---|
| Important spoken-meaning mistakes | ElevenLabs made 121% more | 64 vs 29 · Familiar made 54.7% fewer |
| Finished dubs with any important meaning problem | ElevenLabs had 56.5% more | 36/111 vs 23/111 |
| Background-sound error | ElevenLabs produced 270% more | 11.92 vs 3.22 dB · Familiar 73.0% lower |
| Laughter/reaction loudness error | ElevenLabs produced 221% more | 3.32 vs 1.04 dB · 18 qualifying clips |
| Laughter/reaction shape correlation | Familiar 28.2% higher | 0.960 vs 0.749 |
| Speaker resemblance | Inconclusive | +3.5% point estimate for Familiar; the paired interval crossed zero |
| Source-clip coverage | Familiar 38/38 · ElevenLabs 37/38 | V2 Alpha rejected the 10.94-second clip |
03.Results: ElevenLabs Dubbing v1 API (the stable product)
The broader study: 418 paired clip-language trials across 38 clips and 11 languages. Both systems completed every requested trial.
| Measure | Result | Raw scores |
|---|---|---|
| Review-flagged spoken-output mistakes | ElevenLabs had 95.5% more | 258 vs 132 · Familiar made 48.8% fewer |
| Background spectral error | ElevenLabs produced 261% more | 12.27 vs 3.40 dB · verified scene regions |
| Laughter/reaction shape correlation | Familiar 243.7% higher | 0.759 vs 0.221 · verified event regions |
| Speaker resemblance | Familiar 27.8% higher | 0.461 vs 0.360 · all 11 languages favored Familiar |
| Sound-effect and beat timing | Familiar 24.3% higher F1 | 0.947 vs 0.762 |
| Predicted naturalness | Familiar 10.2% higher | 2.51 vs 2.28, exploratory |
The voice result deserves the emphasis: on the stable API comparison, the dubbed voice measured closer to the real speaker in every one of the 11 target languages, and predicted naturalness was higher too. Familiar generates the voice and the face together, so the win is not either-or: the output that beat the audio-only pipeline on voice also ships the re-performed face.
04.Workflow limits we hit along the way
Quality aside, running a catalogue through ElevenLabs Dubbing is slow by construction. Three product limits, observed August 5, 2026:
- Clips under 11 seconds are rejected. The 10.94-second clip in our set could not be dubbed at all.
- 20 items per batch. Four videos into six languages is already 24 outputs, more than one batch holds.
- 30 concurrent requests. Ten videos into ten languages is 100 requests: at least four waves of manual batching and waiting before generation time even starts.
For a creator dubbing a back catalogue, or even four videos into more than five languages, that turns one job into hours of babysitting. Familiar accepts any clip length, and one upload fans out to every selected language as separate render jobs.
05.Price
| Option | Per minute | What it covers |
|---|---|---|
| Familiar | $2.50 | Translation, dubbed voice, music/effects/ambience preserved, and the face re-rendered |
| ElevenLabs Dubbing | $3 (checkout quote, Aug 5, 2026) | Audio only |
Cheaper, and the $2.50 includes the face. A 60-minute video in one language is $150 complete; published human dubbing packages run $39 to $94 per finished minute for audio alone. The full price study, with sources, is in How Much Does Dubbing Cost in 2026?
06.Translation text vs professional translators
A separate study compared translation text itself against professional work on WMT24++ passages: Familiar's faithful translations reached 96% of a first-pass professional translator's reference-based COMET score (95% CI: 95.2 to 96.9%), under a reference design that structurally favors the professional baseline. Per language, Hindi scored above the professional first pass, French tied it, and Indonesian reached 99.6%.
07.Limitations
- The meaning review is model-assisted and single-pass; independent native double review is pending.
- V2 Alpha speaker resemblance and naturalness were inconclusive at this sample size; the v1 identity result is not pooled with them.
- These are curated real-world clips, not a probability sample of all video. Product limits and prices were observed on August 5, 2026 and may change.
REFERENCES
Dubbing is finally good. See the measurements, then try it on your own video.
