Google Patents a Way to Send Where Sounds Come From During a Call
When you call someone, all sense of where each voice is coming from gets flattened into a single stream of sound. Google is patenting a way to send the position of sounds along with the audio itself, so the person on the other end hears voices and noises where they actually were.
What Google's spatial audio calling system actually does
Ever been on a group video call where everyone's voice seems to come from the exact same place, no matter how many people are talking? That flat, crowded sound is a direct result of how audio is captured and transmitted today: position gets stripped out entirely.
Google's patent describes a system where a device with multiple microphones (think a speaker bar, a phone, or a smart display) figures out the direction each sound is coming from, then packages that location data alongside the audio before sending it over a call. The receiving device, if it has multiple speakers, uses that information to play sounds back in roughly the positions they were originally captured from.
In plain terms: if someone is speaking to your left on the sending end, their voice comes out of the left side of the receiving device. You hear the room, not just the voices.
The first device may capture audio signals in an environment through two or more microphones. The first device may encode the captured audio with direction information. The first device may transmit the encoded audio via the communication link to the second device.
Translation: One device records sound with multiple mics, adds directional data, and sends it to another device.
How direction data is encoded, sent, and rebuilt at playback
The system relies on two paired capabilities: a microphone array (two or more microphones placed apart from each other) on the sending device, and a speaker array on the receiving device.
The sending device uses the slight time and volume differences between microphones to estimate the direction of incoming sounds, a process called spatial audio capture (essentially measuring which mic hears a sound first and how loudly to triangulate position). That direction metadata is then encoded directly into the audio stream before transmission.
On the receiving end, the device decodes the audio and uses the embedded direction data to route sound to the appropriate speakers. The goal is what the patent calls recreating positions, meaning a listener hears a voice from the left if it was originally to the speaker's left, or from across the room if it was farther away.
The patent covers the full pipeline:
- Multi-microphone capture with direction estimation
- Encoding direction data into the transmitted audio stream
- Decoding on the receiving device
- Speaker-array playback that maps decoded positions to physical speaker channels
The system is described as device-agnostic, potentially applicable to phones, smart speakers, or video-calling hardware.
What this could mean for video calls and smart speakers
Right now, most audio calls sound the same whether one person or six are talking: a flat mono or basic stereo mix with no sense of space. Google's run of spatial-audio and hardware-audio filings points toward building this into the kind of shared-space devices, like Nest speakers and Meet hardware, where it would matter most. If a conference room bar can send spatial audio to a remote listener's speaker array, large group calls become meaningfully easier to follow.
For everyday users, the practical gain is reducing the cognitive load of untangling overlapping voices. Spatial separation is one of the oldest tricks the human auditory system uses to tell speakers apart. Putting that back into calls is a small-sounding change that could make long meetings noticeably less exhausting.
Google's sixth filing in the Audio patents we cover since May follows earlier work like sending head-movement data over headphone links and a context-aware smart button.
Sending direction information alongside every audio signal means more data traveling across the call, and if the receiving device has only a single speaker, that extra information does nothing useful. The feature collapses into an ordinary call, which means the experience depends entirely on what hardware both people happen to own.
The deeper cost is environmental. A cluster of microphones can estimate where a voice is coming from in a quiet room, but that estimate becomes unreliable the moment three people are talking at once or someone is moving around. The design solves a clean problem that real calls rarely present.
When both devices have the right speakers and the room cooperates, the added encoding burden is small and the clarity improvement is meaningful. That is a narrow set of conditions, and a feature that performs well only in its easiest scenario carries limited practical weight.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
6 drawing sheets from US 2026/0261813 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →