ALPHA

FAMILIAR VS ELEVENLABS · TWO PAIRED STUDIES · AUGUST 2026

Tested more accurate than ElevenLabs.

Even just on voice dubbing, Familiar is far superior.

The same clips through both systems, only the finished dubs judged. Everything on this page is published: data, confidence intervals, audio.

01 · THE WORDS +121%

ElevenLabs made 121% more important meaning mistakestheir newest V2 Alpha · 111 paired dubs · 3 languages

ElevenLabs64
Familiar29
02 · THE SCENE +270%

ElevenLabs damaged the background sound 270% moremusic, effects, and room ambience · same clips

ElevenLabs11.92 dB
Familiar3.22 dB
03 · THE LAUGHS +221%

ElevenLabs got laughs and reactions 221% more wrong in loudnessshape correlation +28.2% for Familiar (0.960 vs 0.749) · 18 clips / 54 comparisons

ElevenLabs3.32 dB
Familiar1.04 dB
04 · THE VOICE +27.8%

Familiar sounds 27.8% more like the real speakertheir stable v1 API · 418 paired dubs · 11 languages

CLOSER IN ALL 11 LANGUAGES
38/38 clips completed — ElevenLabs rejected the one under 11 seconds $2.50 per finished minute with the face, vs their $3 audio-only checkout quote

ELEVENLABS WEBSITE V2 ALPHA · THEIR NEWEST PRODUCT

Coverage and workflow failures are part of the result.

Same 38-clip set for both systems. Familiar completed 38/38; ElevenLabs Dubbing V2 Alpha rejected the 10.94-second clip under its 11-second minimum — counted as a coverage failure, never a fabricated quality score.

PRODUCT COVERAGE Familiar 38/38 · ElevenLabs V2 Alpha 37/38
COST-CONTROLLED THREE-LANGUAGE SUBSET Mandarin Chinese · Spanish · Japanese

ElevenLabs documents pricing by source duration and target-language count. Using the $3 per dubbed minute per language checkout quote observed on August 5, 2026, the 37 eligible clips (about 36.1 minutes) cost ≈$325 for three languages; all 11 languages (407 outputs ElevenLabs could accept) would cost ≈$1,191. Familiar is $2.50 per finished minute per target language, including audio and visual generation.

V2 ALPHA OUTPUTS Mandarin · Spanish · Japanese complete

All 111 same-clip/language comparisons—222 final audio files—were matched and transcribed. Objective scoring and manual, model-assisted meaning review are complete.

SHORT CLIPS REJECTED10.94-second Jimmy Yang clip cannot be dubbed.

Familiar completed it; ElevenLabs V2 Alpha rejected it under the 11-second minimum.

ITEMS PER BATCHA catalogue job must be split almost immediately.

Just 4 videos into 6 languages is 24 outputs—already too many for one batch.

CONCURRENT REQUESTSMulti-video, multi-language work waits in waves.

Ten videos into ten languages creates 100 requests: at least four waves before generation time is counted.

OBSERVED CATALOGUE WORKFLOWSlow and tedious at creator-catalogue scale.

The Spanish pass alone was slow; larger catalogues can require hours of manual batching and waiting.

Selection fairness. Mandarin, Spanish, and Japanese are a broad, cost-controlled subset. Japanese is a large strategic language; Korean—our team’s strongest native-review language—was not selected. The subset was not chosen simply to maximize Familiar’s score.

111 PAIRED FINAL-AUDIO OUTPUTS

Final-audio results: Familiar vs ElevenLabs V2 Alpha

Same 37 accepted clips in Mandarin, Spanish, and Japanese. Full-lane measures use 111 paired comparisons; laughter/reaction measures use 54 paired comparisons from 18 qualifying clips. Each source clip receives equal weight; 95% intervals resample source clips.

ElevenLabs: 270% more background-sound error 11.92 vs 3.22 dB MAE · Familiar 73.0% lower · paired 95% CI excludes zero
ElevenLabs: 221% more laughter/reaction loudness error 3.32 vs 1.04 dB · Familiar 68.8% lower · 18 clips / 54 comparisons
28.2% higher laughter/reaction shape correlation 0.960 vs 0.749 · 18 clips / 54 paired comparisons · paired 95% CI excludes zero
ElevenLabs: 121% more important spoken-meaning mistakes 64 vs 29 · Familiar 54.7% lower · 111 paired final-audio comparisons
ElevenLabs: 56.5% more finished dubs with any important meaning problem 36 vs 23 · Familiar 36.1% lower · paired 95% CI excludes zero
+3.5% point estimate speaker-resemblance median 0.485 vs 0.469 SIM-o · sample was not large enough to prove a real lead
No clear winner predicted voice naturalness 2.490 vs 2.494 UTMOS · uncertainty is broad
38/38 vs 37/38 source-clip coverage ElevenLabs rejected the 10.94-second clip; no fake quality score was assigned.

How meaning was checked: each finished dub was transcribed with Scribe v2, then manually reviewed against the same approved target meaning. Valid paraphrases were accepted. Tiny background reactions were recorded separately and did not count as important main-speech failures.

Observed in the ElevenLabs website V2 Alpha workflow on August 5, 2026. Product limits may change.

TRANSLATION QUALITY · TWO COMPLEMENTARY STUDIES

Did the meaning make it through?

Final dubbed outputs are tested in real-world clips, then translation text is compared with first-pass professional work against expert post-edits. Claims appear only when the reported confidence interval supports them.

Loading translation-quality results

Checking both final-output and professional-reference evidence…

VALUE · ONE FINISHED-MINUTE PRICE

A complete dub for $2.50 per minute.

Per finished minute, per target language. Translation, dubbed voice, preserved scene audio, and the face re-rendered: included, not sold as separate steps.

Familiar complete dub $2.50 per finished minute
per target language
  • Translation
  • Dubbed voice
  • Music, effects, and ambience preserved
  • The face re-rendered to the new language
60-minute video$150.00

THE SIMPLEST COMPARISON

About 96% less than the midpoint of published human dubbing packages.

Familiar$2.50per minute
Human dubbing midpoint~$60per minute

The published midpoint is about 24× Familiar’s price. This is a price comparison, not a claim that every service has the same production process or quality level.

Published price context in US dollars, per target language
OptionPer finished minuteFor 60 minutesFamiliar price difference
Familiar complete dub$2.50$150.00Translation, voice, scene audio, and the face included
Published human translation + review$10 median
$11.95 mean
$600 median
$717 mean
75.0–79.1% lessThose rates are 4.0–4.8× Familiar’s price
Published all-human dubbing packages$39–$94
~$60 midpoint
$2,340–$5,640
~$3,600 midpoint
93.6–97.3% lessThe midpoint is 24× Familiar’s price
Premium studio dubbingFrom $128From $7,680At least 98.0% lessThe starting rate is 51.2× Familiar’s price
Sources, calculations, and caveatsOpen the pricing notes

The $10 median and $11.95 mean were calculated across the benchmark’s 11 target languages from Netflix’s published rate card. Published all-human package prices come from Voquent, and the premium studio starting price comes from Gotham Lab.

Interpreter context comes from the ATA 2025 Compensation Survey. Sixty-minute totals multiply each per-minute price by 60. Percentage savings use (comparison price − $2.50) ÷ comparison price.

Rates vary by language, talent, casting, studio, revisions, turnaround, and volume. Figures exclude taxes and discounts. The comparisons describe published prices; they do not assume identical workflows or guarantee equivalent quality.

ELEVENLABS DUBBING V1 API · STABLE API COMPARISON · 418 PAIRED TRIALS

Final-audio results: Familiar vs ElevenLabs Dubbing v1 API

This broader stable-v1 API study covers 38 clips and 11 languages. Both systems completed all 418 requested clip-language trials. It is shown beside the newer V2 Alpha study, but the two cohorts are never pooled.

Higher-is-better measures lead with Familiar's gain. Lower-is-better measures lead with how much more error ElevenLabs produced; the raw scores remain directly underneath.

More recognizable voices

Source-speaker resemblance was higher in every target language.

Fewer problems in the final spoken audio

Review found 132 individual issues in Familiar outputs and 258 in ElevenLabs outputs.

The scene survived the dub

Laughter shapes, soundtrack texture, and transient timing stayed closer to the source.

ELEVENLABS DUBBING V1 API · 11-LANGUAGE BREAKDOWN · CLICK TO LISTEN

Hear the difference on ElevenLabs’ stable v1 API

Each section pairs a language-level chart with three short, synchronized listening examples. Choose Original, Familiar, or ElevenLabs; every player starts and stops on the same marked span. The clips illustrate what each metric family sounds like; they do not prove or contribute to the aggregate score.

Loading paired benchmark records and listening examples…

PLAIN-LANGUAGE METRIC GUIDE

What every term means

No acronym should stand between a listener and the result. Open any definition for the short version and the important caveat.

HOW TO READ THIS

One benchmark, several views of quality

This is a paired black-box comparison: the same English source clip and target language were sent to each system, then their final dubbed outputs were evaluated. Higher is better unless a chart explicitly says “lower is better.”

Automated metrics are useful lenses, not a single universal score. Speaker similarity does not measure translation meaning; literal text matching can punish valid paraphrases; acoustic correspondence does not replace listening. That is why each result is shown alongside its definition and source-aligned audio.

Evidence status: complete paired automated audio and transcript diagnostics, plus a model-assisted single-pass review of flagged translation problems. Independent native-language double review and blinded human event adjudication remain pending. No composite score is reported.