Follow
Fix it in post

Sound

AI Dubbing in 2026: Six Million Daily Viewers on YouTube, and Where Dubbed Audio Still Breaks

YouTube auto-dubs by default and ElevenLabs bills per language. When AI dubbing is good enough, when it is not, and a QC checklist for editors.

YouTube now dubs videos by default, and more than six million people a day watch at least ten minutes of the results. For editors, the real question is not whether AI dubbing works. It is which parts of a programme it can be trusted with. Our position: AI dubbing plus a native-speaker review is good enough for explainers, corporate work and most creator content. Leaving it to YouTube without review is a gamble, and for drama, comedy or music-led work it is still the wrong tool.

What YouTube does now, by default

On 4 February 2026, YouTube said auto dubbing was available to everyone in 27 languages, with Expressive Speech, which tries to carry over tone and emotion, in eight: English, French, German, Hindi, Indonesian, Italian, Portuguese and Spanish. It also said that in December it "averaged more than 6 million daily viewers who watched at least 10 minutes of auto dubbed content".

YouTube's help page says the feature "is enabled by default for eligible creators". Videos are skipped if they run over 120 minutes, contain little speech or only music, have speech too fast to dub without an "unlistenable, sped-up" result, or contain copyrighted material. Creators can turn on manual review to preview dubs before release, publish or delete individual tracks, upload their own dubs, or switch the feature off under Settings > Content. YouTube says dubbing has no negative effect on the original video's discovery.

YouTube is also testing lip sync. When Android Authority reported the pilot in October 2025, it handled only 1080p, not 4K, and five languages. That may have changed since.

A dubbing studio. Human dubs still win on timing and performance.

What the paid tools offer

The most visible paid option is ElevenLabs. Its Dubbing v2 works speech to speech and "conditions directly on the original performance", with a "sync-aware" translation step that aligns starts, stops and pacing. It covers more than 90 languages. It produces audio only. As one developer write-up notes, visual lip sync is a separate job in your edit.

On price, ElevenLabs' API pricing page lists Dubbing v2 at $2.20 a minute, and v1 at $0.33 with a watermark or $0.50 without. Its help centre says cost depends on the model, the duration "and the number of languages you're dubbing into". So a 10-minute video in three languages is billed as 30 minutes of dubbing. At the v2 rate that comes to about $66. That is our arithmetic, assuming the cost scales linearly with languages.

Not every vendor charges that way. A comparison published by Perso, a competitor that ranks itself first, says most tools bill once per source minute regardless of how many languages you add. It lists HeyGen lip-synced dubbing at about $0.24 a minute and Rask AI at about $2.40. Treat those as vendor claims and check the current plan before quoting.

A voice-over booth with a broadcast microphone.

Where dubbed audio still breaks

The most detailed public test we found is again a competitor's. Familiar Labs compared its product with ElevenLabs and reported 64 important translation errors from ElevenLabs against its own 29, and "270% more background-sound error: the laughter, the music, the ambience". The study is self-interested and the numbers should not be quoted as neutral. But the failure categories match what anyone who has QC'd a dub would check, and YouTube's own exclusions point the same way: music, fast speech, little dialogue.

In our experience the breaks fall into a few predictable places:

  • Names and terms. Brand names, people and product terms get translated, mispronounced or both.
  • Numbers. Figures, units, currencies and dates are where a mistake turns into a factual error and, in corporate work, a legal one.
  • Non-speech vocals. Laughter, breaths, crosstalk and overlapping speakers get dropped, doubled or turned into speech.
  • The music bed. When the tool separates voice from music itself, the bed can pump or lose detail under the new dialogue.
  • Timing. Languages that run longer than the source get squeezed, and the dub drifts against the cuts.
  • On-screen text. The voice now speaks Italian while the lower thirds and graphics remain in English.

A decision rule

  1. Leave it to YouTube only for low-stakes creator content, and only with manual review switched on, so nothing goes live unheard.
  2. Deliver your own AI dub plus a native-speaker review for corporate films, product explainers and anything with claims, prices or brand names. Upload it as your own dub track so you control the version.
  3. Use human dubbing or subtitles for drama, comedy and music-led work, where the performance is the product.

Budget the review, not the generation. If a native speaker needs about an hour per language for a 10-minute film (our working figure, not a published one), the reviewer costs more than the dubbing fee. The review is also the part that protects the client.

The QC checklist

  • Give the tool a clean dialogue stem and a separate music-and-effects mix, not the final mix, so it does not have to separate them itself.
  • Supply a glossary of names, brand terms and anything that must not be translated.
  • Check every number, unit, price and date against the script.
  • Listen for laughter, breaths and crosstalk at every point where the source has them.
  • Check that the music ducks under the new dialogue the way it did under the original.
  • Watch for sync drift at cut points, especially in languages that run long.
  • List on-screen text that needs a localised graphic or a subtitle.
  • Match loudness to the original track before upload.

As of 2 October 2026 these tools are changing month to month. Recheck prices and language lists before you quote a job.

Keep reading