Disney Patents a System That Resyncs Body Language When It Dubs Video Into Another Language
Bad dubbing is one of those things you notice immediately: the words don't match the mouth, and the hand gesture lands a beat after the sentence it belongs to. Disney has filed a patent for a system that tries to fix both problems at once.
How Disney's dubbing patent retimes gestures automatically
Every time a movie studio dubs a film into another language, a familiar problem shows up: the actor's body language was designed for the original words, not the translation. A gesture that punctuates the end of an English phrase might come half a second too early or too late when the French version runs longer or shorter.
Disney's new filing describes a system that breaks a video into its audio and visual parts, translates the speech, and then figures out how to line up the actor's physical gestures with the new translated version. It can either shift the timing of a gesture already in the video, or adjust the pacing of the translated audio so the two click back into sync.
The patent also covers lip movements, so the goal is a dubbed clip where the mouth, hands, and voice all match the new language without a human editor painstakingly adjusting each one.
modifying a temporal location of the first gesture in the multimedia asset or modifying a timing of a translated audio segment to set a second temporal alignment between the first gesture and a translation of the utterance in the translated audio segment.
Translation: It fixes the mismatch by shifting either the body movements or the translated speech so they line up again.
How the system aligns gestures to translated speech timing
The system starts by pulling a video apart: the audio track is extracted separately from the visual frames. The audio goes through a translation process to produce a new spoken version in the target language.
At the same time, the video frames are analyzed to identify body gestures and the exact moment each one occurs. The system then maps out what it calls a temporal alignment (basically a timestamp relationship) between a spoken utterance and the gesture that accompanied it in the original.
Once the translated audio is ready, the system compares its timing to the original. Translations rarely run the same length as the source, so the alignment is off. The patent describes two ways to fix this:
- Move the gesture to a new position in the video timeline to match when the translated phrase lands
- Adjust the pacing of the translated audio clip so it hits its end at the same moment the original gesture does
Lip movements get the same treatment, generating new lip-sync that corresponds to the translated words. The output is a synchronized multimedia asset, a single video file where gestures, lips, and translated audio are all timed to each other.
… identify body gestures of individuals in the video segment, generate translated body gestures, and generate translated lip movements.
Translation: The system detects physical gestures and creates matching visual movements along with the new voiceover.
What this means for dubbed movies and global streaming
For viewers, the difference between good and bad dubbing is almost physical: a gesture that doesn't match the speech pulls you out of the story immediately. If this system works in practice, it could raise the floor on dubbed content across Disney's streaming catalog, which includes films in dozens of languages.
Disney keeps filing on AI-driven media production tools, and this one sits at the intersection of translation technology and video editing automation. The practical path to shipping it requires solid AI-generated gesture detection and body-pose manipulation, which are active areas of computer-vision research. None of that new hardware is needed, but the software pipeline has to be reliable enough that it doesn't produce results more distracting than the original sync problem.
This is the sixth Disney filing we've tracked in our AI vision coverage since July, following one on pinning masks to objects and one on tracking live show scenes.
The core pieces this system needs already exist as everyday software tools: automatic translation, body-movement detection, and video editing. What Disney is describing is a specific way to connect them into a single pipeline that takes a finished film, translates the dialogue, and adjusts the actors' mouth shapes and hand gestures to match the new language.
The hard part is the gesture step. Deciding which arm movement belongs to which word is tricky when actors are in constant motion, and a wrong guess would be immediately obvious to any viewer watching at home.
The patent commits to specific outputs and specific methods rather than staying vague, which means an engineering team could build toward it without new hardware or waiting on a scientific breakthrough. The shortest path to a product is refining the matching logic until it clears a quality bar Disney would actually release.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
5 drawing sheets from US 2026/0281513 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →