New Google Patents · Filed Jun 2, 2025 · Published Jul 16, 2026 · verified — real USPTO data

Patent Turns a Single Audio Track Into Surrounding Sound From Every Direction

Google is patenting a way to take a single-channel audio recording and use a generative AI model to reconstruct a full, three-dimensional sound field around it. Think of it as the audio equivalent of turning a flat photo into a 3D scene.

Google Patent: AI-Generated Spatial Audio From Mono Signals — figure from US 2026/0205753 A1
Figure from the official USPTO publication.
Publication number US 2026/0205753 A1
Applicant GOOGLE LLC
Filing date Jun 2, 2025
Publication date Jul 16, 2026
Inventors Willem Bastiaan Kleijn, Michael Chinen
CPC classification 381/1
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (May 20, 2026)
Parent application is a National Stage Entry of PCTUS2022081834 (filed 2022-12-16)
Document 25 claims

How Google turns one audio track into surround sound

Imagine you recorded a live concert on a single microphone. Everything you captured is flat, with no sense of which sounds came from the left, the right, or behind you. Getting that full spatial experience normally requires recording with many microphones at once.

Google's patent describes a system that can take that single recording and intelligently generate the missing spatial information around it. Instead of hard-coding fixed rules for where sounds should appear in space, it uses a generative AI model that has learned from training data how sounds plausibly exist in a real environment. It then uses that learned knowledge to place sounds in a believable three-dimensional arrangement.

The system can work in two modes: it can try to reconstruct what the original spatial recording actually sounded like, or it can simply generate a convincing spatial version on its own, guided by what it has learned. Either way, you end up with multi-channel audio from a source that started as just one track.

How the generative model samples spatial placement data

The patent centers on something called a Spatial Audio (SA) model, which describes sound in terms of individual audio sources (called pseudosources) and their positions in space. Traditional systems encode that spatial position information directly and deterministically, meaning they store exactly where each sound is placed. Google's approach replaces that fixed encoding with a stochastic (probability-based) generative process.

Specifically, the system uses a conditional model distribution, which is a learned probabilistic model that takes two inputs:

  • A conditioning sequence derived from the single-channel (mono) reference audio signal
  • A conditioning sequence that describes what kind of spatial arrangement is desired

From those two inputs, the model samples a spatial arrangement sequence rather than looking one up. Sampling here means the model draws a plausible answer from a range of statistically likely options, much like how a language model generates text word by word. The sampled spatial arrangement is then combined with the original audio to produce a full multi-channel output.

The approach also allows the system to lean entirely on what it learned during training, without needing explicit spatial instructions. That makes it flexible enough to work even when good reference spatial data isn't available.

What this means for streaming, calls, and VR audio

For consumers, this kind of technology could improve spatial audio quality in situations where full multi-microphone recordings were never made, such as older music archives, phone calls, or podcast recordings. If Google integrates this into products like YouTube, Google Meet, or its Immersive View features, your existing mono or stereo content could be presented in a richer, more three-dimensional way without anyone re-recording anything.

For the industry more broadly, replacing deterministic spatial encoding with a generative model is a meaningful architectural shift. It means the system can fill in missing spatial information intelligently rather than defaulting to silence or a fixed stereo spread, which matters most for VR, AR, and next-generation audio codecs where spatial accuracy is central to the experience.

Editorial take

This is a genuinely interesting signal from Google's audio research group. Spatial audio has been a quiet battleground across Apple, Sony, and Dolby for years, and using a generative model to synthesize spatial information rather than encode it rigidly is a real departure from how the field typically works. It won't make headlines the way a new phone does, but it's the kind of foundational codec-level work that could improve audio quality across a lot of Google's products.

Which company should we read for you?

We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.

Get one Big Tech patent every Sunday

Plain English, intelligent commentary, no hype. Free.

Source. Full patent text and figures from the official USPTO publication PDF.

Editorial commentary on a publicly published patent application. Not legal advice.