Disney Patents an AI System That Scores How Closely Two Voice Performances Match
Judging whether one voice actor sounds like another is normally a gut call. Disney is filing a patent to turn that judgment into a number.
How Disney's voice-matching score actually works
Every time a studio records a replacement voice line, a director has to decide whether the new take matches the original delivery well enough. That call is entirely subjective, and in a production with hundreds of lines, it takes a lot of time.
Disney's patent describes an AI tool that listens to two speech samples and produces a similarity score that captures things like rhythm, pace, and pitch pattern, what linguists call prosody. Feed it the original and the re-recorded take, and it tells you how close they are.
That score can then trigger an action. The system might flag a take for human review, automatically accept it, or route it back to the voice actor. The goal is to give production teams a consistent, repeatable standard instead of relying entirely on a director's ear.
inputting a first speech sample into a prosody encoder, wherein the prosody encoder is trained as part of a text-to-speech model; inputting a second speech sample into the prosody encoder; generating, using the prosody encoder, a first representation of the first speech sample and a second representation of the second speech sample; …
Translation: The system feeds two different voice clips into a specialized neural network to create digital profiles of how they sound.
Inside Disney's prosody encoder and comparison pipeline
The system centers on a component called a prosody encoder, a neural network trained as part of a text-to-speech (TTS) pipeline. Text-to-speech models learn to generate natural-sounding speech, and in doing so they develop an internal representation of what makes speech sound a particular way rhythmically and tonally. Disney is repurposing that learned knowledge as a measurement tool rather than a generation tool.
Here is how the process runs:
- A first speech sample (say, the original voice actor's line) is fed into the encoder.
- A second speech sample (say, a dubbing actor's version) is fed in separately.
- The encoder converts each sample into a vector representation, a compact numerical fingerprint of that sample's prosodic qualities.
- The system compares the two vectors to produce a metric value, a single number indicating how similar the two performances are.
The claim is intentionally broad about what happens next. Based on that metric value, the system determines an action, which could mean flagging the result for a human, automatically approving a take, or logging the score for quality-control records.
The key design choice is using an encoder trained inside a TTS model. Because TTS models must learn nuanced speech characteristics to generate convincing audio, the encoder has already learned a rich internal language for describing prosody. Disney is borrowing that learned language for comparison, rather than training a separate scoring model from scratch.
The method compares the first representation and the second representation to determine a metric value and determines an action to perform based on the metric value.
Translation: It then compares those digital profiles to calculate a similarity score that dictates the next step in the process.
What this means for voice casting and dubbing at scale
Voice dubbing and localization are expensive, high-volume operations for a company like Disney. A single animated feature can require dozens of language versions, and each one needs to match the emotional cadence of the original. Right now, consistency depends heavily on individual directors and recording engineers, which means standards vary across sessions, studios, and languages. A quantitative similarity score would give producers a shared benchmark.
The system also has obvious applications in automated quality-control pipelines, where a computer could do a first pass on thousands of recorded lines before a human ever listens. The broader picture for AI in entertainment production is moving fast, and new Big Tech patents in speech and voice technology from studios and platform companies are increasingly treating audio quality as something that can be measured and automated, not just heard.
That makes this Disney's second filing in Voice & speech AI we've tracked since June, following rewriting tour scripts in any voice.
Matching a character's voice across a film, a dubbed translation, and a video game expansion is expensive guesswork. Every re-recorded line currently gets approved or rejected by whoever is in the room that day, and that person's judgment shifts with fatigue, deadlines, and personal taste. For studios running dozens of productions at once, that inconsistency adds up into real money and audible mistakes.
Disney's patent addresses this by converting that gut-check into a repeatable score, using a system already trained to understand what makes a voice sound like itself. The document leaves open the question of what score is good enough, which is where the hard work actually lives.
But establishing a consistent measurement at all is the necessary first step, and the problem it targets is real enough to justify the effort.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
6 drawing sheets from US 2026/0253579 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →