Sony · Filed Aug 20, 2025 · Published Sep 3, 2026 · verified — real USPTO data

Sony Patent Makes Replacement Voice Recordings Sound Like They Belong in the Scene

Dubbing a film or show is hard enough, but replacing a voice while keeping the room's sound is one of audio's messiest unsolved problems. Sony is filing a patent for an AI that handles it automatically.

A system processes an input surround sound signal to extract reverberation characteristics and then applies them to a foreign-language voice signal. Drawing from patent filing US 2026/0261817 A1.
A system processes an input surround sound signal to extract reverberation characteristics and then applies them to a foreign-language voice signal.
See all 8 drawings from this filing ↓
Publication number US 2026/0261817 A1
Applicant Sony Group Corporation
Filing date Aug 20, 2025
Publication date Sep 3, 2026
Inventors Akira TAKAHASHI
CPC classification 704/225
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (May 29, 2026)
Parent application is a National Stage Entry of PCTJP2024003523 (filed 2024-02-02)
Document 12 claims

What Sony's auto-reverb dubbing system actually does

Imagine watching a foreign film where the villain speaks in a big cathedral, and the dubbed English voice sounds like it was recorded in a closet. That jarring mismatch is one of the most common complaints about dubbed content, and fixing it by hand takes serious studio time.

Sony's patent describes a system that listens to the original audio track and automatically figures out two things: what the voice sounds like on its own, and what the room around it sounds like. Then it applies that same room character to the replacement dubbed voice, so your ears don't notice the switch.

The key is a trained AI model that can separate those two layers without needing a separate clean recording. It works from the mixed audio that already exists in the finished film or show.

From the filing · CLAIM 1
1 . An information processing device comprising a reverberation extraction unit that receives an input of an acoustic signal, and separates and extracts a 1ch impulse response corresponding to a direct-wave component and an impulse response of a reverberation component.

Translation: The system breaks down incoming audio to isolate the direct voice from the background echo.

How the AI pulls room echo apart from a direct voice signal

The system takes in a multi-channel audio signal (think 5.1 or similar surround-sound tracks that most films are mastered in) and passes it through a trained neural network.

That network is designed to estimate two things simultaneously:

  • A direct-wave impulse response, a mathematical fingerprint of the original voice as it would sound with no room influence at all (called a 1-channel, or 1ch, impulse response)
  • A reverberation impulse response, a fingerprint of the acoustic echo and room tone that the microphone captured, spread across the surround channels

An impulse response is essentially a recording of how a room reacts to a single sharp sound, it captures all the echoes, reflections, and decay that make a space sound like a cave versus a conference room.

Once the system has both fingerprints, it can apply the room's impulse response to the new dubbed voice. The replacement voice gets wrapped in the same acoustic character as the original, so it sounds like it was recorded in the same space. The neural network was pre-trained on multi-channel audio pairs to learn the patterns that distinguish dry voice from room reflection.

From the filing · THE ABSTRACT
An information processing device that automatically adds an effect such as reverberation similar to an effect in original content to a dubbed-in voice is provided.

Translation: This tool automatically gives replacement dialogue the exact same room echo as the original recording.

What this means for foreign film and game localization audio

For localization studios, dubbing is expensive partly because audio engineers spend hours manually adding reverb, eq, and room matching to replacement voice tracks. A system that extracts room character automatically from the existing audio could cut that work down significantly, and it could run on content that never had a clean stem (a separate dry vocal track) preserved.

For viewers, the practical result is dubbed content where voices feel like they belong in the scene. Sony's pattern of audio-processing AI filings suggests the company is building toward tools that serve both its entertainment production pipeline and its professional audio equipment lines, though nothing announced yet connects directly to this patent.

Sony's third filing in the voice and speech AI patents we've tracked since May builds on earlier work around converting whispers to normal speech and voice-based game controls.

Editorial take

The ship path here is reasonably short compared to most AI audio patents. The core requirement is a trained model and a software pipeline, no new microphone hardware, no changes to how films are mastered, and no cooperation needed from the original studio beyond access to the mixed audio track.

What has to exist first is a large enough dataset of multi-channel recordings with known room properties to train the model properly. That is a real constraint, but Sony runs some of the world's largest music and film production operations, which means it likely has access to exactly that kind of archive.

The remaining question is accuracy at the edges: unusual acoustic spaces, non-standard room shapes, or heavy post-processing in the original could all confuse the separation. The patent does not describe how the system handles those cases, which suggests the hard part of this work is still ahead.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

8 drawing sheets from US 2026/0261817 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.