Google Patents an AI System That Rebuilds Your Voice from Head Vibrations to Cut Out Background Noise
What if your headset could hear you through the noise by feeling how your head vibrates when you talk? That's exactly what Google has patented: a system that rebuilds your voice from bone vibrations, so loud environments can't ruin your call.
How Google's headset separates your voice from any noise around you
You're on a call in a crowded coffee shop, a busy street, or a noisy office. The person on the other end can barely make out what you're saying, and no amount of speaker-volume cranking fixes it because the problem is on your end.
Google's new patent describes a way around that. A head-worn device (think earbuds or a headset with a built-in sensor) would pick up two things at once: the ordinary microphone audio, which includes your voice and all that background noise, plus a separate reading from a vibration sensor pressed against your head. When you speak, your skull vibrates in a pattern unique to your words. That sensor captures those vibrations directly, so the surrounding din simply doesn't show up in the reading.
An AI model then uses both signals together, alongside a stored voice profile built from samples of how you sound, to reconstruct a clean version of what you said, in your actual voice. The result is sent to the other person instead of the noisy original.
… receiving, from a vibration sensor in contact with a head of the user, an inertial signal that represents vibrations of the head induced by the utterance …
Translation: A sensor touching your head picks up bone and tissue vibrations when you speak.
How the ML model fuses three signals into a clean voice waveform
The system draws on three separate inputs to produce a clean audio output:
- Microphone audio: a standard recording of your speech mixed with ambient noise, captured as usual.
- Inertial signal: readings from a vibration sensor (sometimes called an IMU or accelerometer) pressed against the surface of your head. Because solid bone conducts vibration differently from air, your speech shows up here even when background noise doesn't.
- Voice embedding: a compact numerical fingerprint (essentially a short list of numbers that describes what your voice sounds like) stored in advance and pulled in at inference time.
A machine learning model takes all three inputs and generates a synthesized waveform, a digitally constructed audio signal that represents what you said, cleaned of ambient noise and shaped to match your voice characteristics.
The term "waveform" just means a digital audio file, the same format as any audio your phone plays. What makes this one unusual is that it isn't simply the original recording with noise subtracted; it's a new audio signal built by the model from the evidence of all three inputs.
The patent doesn't specify a single hardware form factor, but the claim that the vibration sensor must be "in contact with a head of the user" points clearly to head-worn devices such as earbuds, over-ear headphones, or dedicated headsets.
… generating, by a machine learning model and based on (i) the audio signal, (ii) the inertial signal, and (iii) the voice embedding, a synthesized waveform that represents the utterance in the voice of the user and independently of the ambient noise …
Translation: An AI combines air sound, head vibrations, and your voice profile to rebuild your speech without background noise.
What this means for calls in loud, busy places
For anyone who regularly takes calls outside or in shared spaces, the practical payoff here is significant. Current noise cancellation works by analyzing incoming audio and trying to identify which parts are speech versus noise, a guessing game that gets harder as environments get louder. A bone-conduction sensor sidesteps that problem because it only picks up vibrations from the wearer's skull, so the guesswork largely disappears.
Google's run of head-worn audio processing filings suggests the company sees wearable audio as a serious long-term area. For the person on the call, though, the concrete change is simpler: the other person would hear you clearly in situations where today they'd ask you to step outside.
Google's tenth filing in our wearable tech coverage since May connects to earlier applications like one on tracking chewing and jaw grinding and one on keeping both earbuds in sync.
The practical promise here is simple: you take a call in a loud place and the person on the other end hears you clearly. That happens because the system leans on a sensor pressed against your head, picking up the physical rumble of your voice through bone rather than fighting to isolate your words from surrounding din.
The voice profile piece introduces a real trade-off. The system rebuilds your voice from scratch using a stored sample, which could leave you sounding oddly pristine or slightly unlike yourself, and that uncanny quality might bother people on the receiving end more than a little background noise ever did.
Still, the failure this prevents is one most people have accepted as unfixable: stepping outside, wind hits, and suddenly you're useless on a call. If this works as described, that problem disappears and you only notice because nobody asks you to repeat yourself.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
6 drawing sheets from US 2026/0279329 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in