Google Patents an AI That Separates Nearby Voices From Background Noise by Distance
Your smart speaker hears you just fine, but it also hears the TV, the dog, and whoever is talking in the next room. Google is training an AI to sort sounds not by who made them, but by how far away they were.
How Google's mic AI tells your voice from the room's noise
A smart speaker sits on your kitchen counter while the television blares in the living room and a conversation drifts in from outside. For you, separating all of that is effortless. For the microphone, every sound arrives jumbled together.
What Google is working on here is an AI that listens to that messy mix and automatically splits it into two buckets: sounds coming from close to the microphone and sounds coming from far away. The idea is that the sounds you actually care about, your voice, for example, are usually nearby, while the noise you want filtered out tends to come from a distance.
To teach the AI this trick, Google generated thousands of simulated audio scenes using a computer program that models how sound behaves in physical spaces. That synthetic training data gave the neural network enough examples to learn the telltale acoustic fingerprints that separate a nearby voice from a distant one, without needing to record thousands of real rooms.
… the synthetic training data having been generated by an acoustic simulator to capture acoustic characteristics that vary based on a respective distance from the respective audio input component …
Translation: The system uses a virtual environment to learn how sound changes as it travels different distances to a microphone.
How the neural network learns near vs. Far acoustics
The patent describes a distance-based sound separation system built on a neural network trained entirely on synthetic audio data.
The core idea is to divide any incoming audio mixture into two categories:
- Near sounds: audio from sources physically close to the microphone
- Far sounds: audio from sources farther away
To create the training data, Google uses an acoustic simulator, a software program that models how sound waves behave in physical environments: how they reflect off walls, how they lose energy over distance, and how they arrive at a microphone from different angles. These simulated scenes generate audio mixes where the near/far ground truth is already known, giving the neural network clear examples to learn from.
The neural network is then trained to pick up on acoustic characteristics (subtle cues like reverberation, or the way a room's echo builds on a sound as it travels farther) to predict which parts of an incoming audio mix belong to the near bucket and which belong to the far bucket.
The trained model is then delivered, ready to run on whatever device or service needs it. The patent covers the full pipeline: generating the training data, training the model, and outputting the finished neural network.
… training, based on the training data, a neural network to predict near sounds and far sounds in a particular audio mixture received by a particular microphone.
Translation: The software is taught to identify and separate audio based on whether the source is close to or far from the device.
What this means for voice assistants and smart speakers
For anyone who uses a voice assistant, a video call app, or a smart speaker, this kind of technology is the difference between a device that understands you and one that keeps asking you to repeat yourself. Right now, most microphone noise-cancellation works by learning to suppress everything that isn't speech. A distance-aware system could be more selective, keeping a nearby whisper while dropping a loud conversation from across the room.
The reliance on synthetic training data is worth noting too. Collecting real-world audio at scale raises serious privacy questions, so building the training pipeline entirely inside a simulator sidesteps that problem. Google's approach here fits a pattern of audio AI development that trades real-world messiness for controlled, labeled simulation data, a trend well covered across Big Tech patent news in the audio processing and on-device AI space.
Training a model entirely on fake, computer-generated sound is faster and more private than recording real people. But fake sound is still fake.
A computer can only guess how noise bounces around a real kitchen or a busy café. Wherever that guess is wrong, the model will struggle.
The bet is reasonable. Acoustic simulation has improved a lot. But the model's true worth will show only in real rooms that no simulator ever thought to recreate.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
17 drawing sheets from US 2026/0245575 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →