You have a finished video in English and a customer base that reads Arabic. The temptation is to run it through a dubbing tool and publish. Before you do: AI video dubbing is not one job. It is three, solved to very different degrees, and only one of them fails where you can see it.
AI video dubbing is three problems, not one
AI video dubbing splits into three independent problems: translating the script, generating a voice that says it, and re-timing the speaker's mouth to match. Each is built by different teams, fails in a different way, and needs a different person to catch the failure. Treating them as one setting in one tool is why most dubs land badly.
| Layer | How far along | How it fails | Who catches it |
|---|---|---|---|
| Translation | Strong for common pairs, weak on register | Right words, wrong voice: a formal brand reads casual, a product name gets translated | A native speaker in your sector |
| Voice generation | Strong for steady narration, weak on emotion | Flat delivery, wrong emphasis, mangled numbers and names | Anyone who hears the whole track |
| Lip-sync | Least settled, by a distance | Shapes that nearly match, soft teeth, drift before each cut | Your viewer, in seconds |
Order matters as much as quality. A translation error propagates into the voice track and then into the mouth shapes, so a wrong script gets rendered flawlessly, three times over. Sign off the text before anything is spoken, and the audio before anything is animated.
That sequencing is the cheapest quality control available, and almost nobody does it.
Why does the lip-sync still look wrong?
Lip-sync fails because human audiovisual perception is far more sensitive than the model is precise. Broadcast engineering measured the tolerance decades ago: viewers begin detecting a mismatch when sound runs ahead of picture by roughly forty-five milliseconds. A generated mouth does not just need the right shapes. It needs them at the right instant, on every syllable, for the whole clip.
That threshold comes from Recommendation ITU-R BT.1359-1, used by broadcasters for years to judge when a feed is out of sync. It is a property of the viewer, not of the pipeline.
An older finding explains why a near-miss is worse than you expect. In 1976, McGurk and MacDonald reported in Nature that dubbing the sound ba onto a mouth articulating ga makes listeners hear a third syllable, da. Vision overrides hearing. An almost-right mouth drags speech perception off course, which is why the answer to a bad dub is so rarely a better voice.
The footage decides how much the model has to work with:
- Facial hair. A beard hides the lip line the model must redraw, and the seam lands in it.
- Angle and distance. Models are strongest head-on; a turned head loses the far corner, a wide shot leaves too few pixels.
- Occlusion and cuts. A hand across the face leaves nothing to key from, and sync clean at the top of a shot drifts by the end.
Watch the source with the sound off before you promise anything
Much of the GCC corporate footage I am handed shows a bearded presenter at a slight angle in a wide two-shot. All three are lip-sync poison, and no budget fixes them afterwards. So I watch thirty seconds muted before quoting, and if the mouth is small, shadowed or turning I say so on the first call.
Should you clone the speaker's voice or use a synthetic one?
Clone the speaker's voice when the person is the brand and the audience already knows how they sound; use a stock synthetic voice when the narrator is interchangeable. The deciding factor is usually not quality. It is consent — a cloned voice needs documented, specific permission from the person who owns it, covering the languages and uses you actually intend.
For plain narration few audiences can tell a good synthetic voice from a good clone, so what a clone buys is identity: a founder or a coach whose voice is part of what the audience came for.
If you do clone, keep a record naming the person, the languages and channels the clone may appear in, and how permission is withdrawn, signed by them rather than their manager. I am a technologist, not a lawyer, and rules on voice and likeness differ sharply between countries, so take the legal question to a lawyer where you publish. What I will say flatly is that a voice cloned from a video you found online is not one you have permission to use.
Arabic: which Arabic are you dubbing into?
Arabic is not one target. Modern Standard Arabic reads as authoritative and travels across every market you sell into; Gulf, Egyptian or Levantine dialect reads as human and local but narrows the audience and can sound faintly wrong in the wrong country. That choice is editorial. It belongs to whoever owns the brand voice, not to a dropdown in a dubbing tool.
The standards bodies treat this as more than an accent. In ISO 639-3, Arabic is a macrolanguage, with separate codes beneath it for Standard, Egyptian and Gulf Arabic among others. When your system says ar and nothing more, it has not made the decision. It has hidden it, and every downstream tool guesses toward whatever its training data was heaviest in, which is rarely Gulf.
My working rule is simple enough to say in a meeting. Government, banking, healthcare and pan-GCC corporate communication get Modern Standard Arabic. So does anything a compliance team reads. The alternative sounds regionally partisan in a room where that matters. Retail, food, fitness, consumer apps and social-first short form get the dialect of the market you are selling in. MSA in a casual retail clip sounds like a news bulletin advertising shawarma. Never mix the two inside one video, which is what happens when three people touch a translation and nobody wrote the rule down.
There is a mechanical wrinkle too. Arabic's pharyngeals and emphatics are produced at the back of the vocal tract and carry almost no lip shape, so a mouth model trained mostly on English has less to key off in Arabic than in French.
When is a subtitle the better answer than a dub?
A subtitle beats a dub whenever the speaker's own voice carries the credibility, whenever the clip will be watched on mute, and whenever nobody on your team can review the dubbed audio. A subtitle error is visible to you and cheap to fix. A dub error is invisible to you and obvious to your audience.
So ask the uncomfortable question: if this dub is bad, who on your side will ever find out? If the honest answer is nobody, you are choosing between a subtitle and an unmonitored risk.
- Testimonials and founder video. Authenticity is the asset, and a synthetic voice spends it.
- Silent-autoplay placements. Those views hear nothing, so on-screen text is the localisation.
- Overlapping speakers. Panels and crosstalk defeat separation and voice.
- No consent, and a recognisable face. A stranger's voice on a known face is worse than text.
- Accessibility. WCAG 2.2 Success Criterion 1.2.2 makes captions for prerecorded synchronised media a Level A requirement, and swapping the audio language produces no caption.
The case where I would not recommend dubbing is the one clients ask for most: a hero brand film, close on the founder's face, into five languages next week with no native reviewer booked. Subtitle it, and spend the budget on the next twenty videos, where the mouths are smaller and the stakes lower.
What a native reviewer has to check before you publish
A native reviewer works from a list, not a feeling: register, terminology, brand and product names, numbers and dates as spoken, the last seconds of every shot for sync drift, and any on-screen text the new audio contradicts. Hand them the source script and the translation side by side.
- Register. Does it sound like your brand talking, or like a translation of your brand talking?
- Terminology and names. Industry terms, products and founders against a fixed glossary and pronunciation lexicon.
- Spoken numbers. Prices, phone numbers, dates. This is where synthetic voices embarrass you.
- Sync drift. Scrub the final second of every shot, not the opening frames.
- On-screen text. Lower thirds still say the old language, and Arabic subtitles carrying Latin brand names or URLs can reorder themselves.
- The mix. Does the dialogue sit in the original acoustic space, or float on top of it?
Ask for the M and E stems before you agree to dub anything
Dubbing has always assumed you hold a music-and-effects track with the dialogue stripped out. Most clients hand me a finished mixdown, so the choice becomes losing the music bed or running source separation and inheriting its artefacts under every line. If the video is recent the stems are usually still on the editor's drive. Ask first.
Dubbing is a pipeline step, not a one-off job
Dubbing gets cheap only when it stops being a task someone performs and becomes a step something runs: one source script per video, a fixed glossary and pronunciation list per language, a render step per locale, and a review queue where an approver sees script and output together.
The payoff is version control. When the English master changes — a price, a claim, a name — every language regenerates from the corrected script, instead of six people re-editing six videos and five forgetting.
There is a shortcut worth knowing. If the source is a presenter reading to camera, dubbing that take is the harder road: regenerating the presenter in the target language sidesteps lip-sync instead of fighting it, which is what an AI avatar video generator pipeline is for. Dubbing earns its place on footage you cannot regenerate — real locations, real customers, a founder on a real stage.
Tech Vision Era builds and operates pipelines like these for clients, handed over as code the client owns. It is a build engagement, not a product, which matters here: the glossary, the dialect rule and the reviewer queue are editorial decisions that belong in your own system.
So here is this week's decision. Take one existing video and one target language, produce a subtitled cut and a dubbed cut, and put both in front of a native speaker who knows your sector, source script beside them. That comparison will tell you more about whether AI video dubbing belongs in your workflow than any tool evaluation.