Google Patents a System That Asks You to Correct Its Own Listening Mistakes
Google's voice assistant sometimes mishears your words as a command, or misses you entirely when you actually said 'Hey Google.' A newly filed patent describes a system that, when it's genuinely unsure, simply asks you what you said and uses your answer to get better over time.
How Google's assistant learns from your corrections
You're across the room when you call out "Hey Google," but nothing happens. Or worse, your assistant springs to life in the middle of a conversation you weren't directing at it. Both problems come down to the same root cause: the AI has to draw a line between "that was definitely a wake word" and "that definitely wasn't," and some audio falls awkwardly in the middle.
Google's patent describes a way to handle that gray area by looping you in. When the system's confidence score lands in an ambiguous middle zone, below its main threshold but above a secondary, lower one, it pauses and asks you directly: "Did you mean to say a hotword?" Your yes or no then nudges the main threshold up or down, calibrating the system to your voice and environment over time.
The effect is a voice assistant that tunes itself to you specifically, rather than relying entirely on a one-size-fits-all model baked in at the factory. Google keeps filing on on-device AI personalization, and this fits that pattern: the learning happens based on real feedback from real users, not just lab recordings.
… prompting the user to indicate whether or not the spoken utterance includes a hotword; receiving, from the user, a response to the prompting; and adjusting the primary threshold based on the response.
Translation: The device asks if it heard the wake word correctly and uses your answer to update its sensitivity.
How the two-threshold system decides when to ask
The patent describes a two-threshold architecture sitting on top of a machine learning model that already scores audio for wake-word likelihood.
How the thresholds work:
- The primary threshold is the normal bar for triggering the assistant. Clear audio of someone saying "Hey Google" clears this easily.
- The secondary threshold is a lower bar. Audio that clears this but not the primary is treated as genuinely ambiguous, not a clean miss and not a clear hit.
- When audio lands in that gap, the system prompts the user (likely a notification or a short audio cue) asking whether they intended to say a hotword.
The user's answer is then used to adjust the primary threshold going forward. If you said yes and the assistant missed you, the bar comes down slightly. If you said no and it nearly triggered incorrectly, the bar moves up.
The patent also notes the model itself can be retrained on the labeled examples your responses create. Over many interactions, both the threshold and the underlying model become more accurate for your specific voice, accent, and typical environment (a noisy kitchen, a quiet office).
Critically, this all hinges on the secondary threshold being set conservatively enough that users aren't bombarded with prompts. The system only asks when the score is close enough to matter.
… determining that the predicted output satisfies a secondary threshold that is less indicative of the one or more hotwords being present in the audio data than is a primary threshold; …
Translation: It notices when speech is close to a wake word but not quite sure enough to trigger normally.
What this means for 'Hey Google' false alarms
False activations are one of the most complained-about features of any voice assistant, and missed activations are just as frustrating. Both stem from the same problem: a global model trained on millions of voices can't perfectly account for the way your voice sounds in your home. The current fix is to retrain on large datasets and hope for statistical improvement. This patent's approach is more targeted: turn every uncertain moment into a labeled data point from the actual user who matters.
For you as a user, it could mean fewer moments where your assistant hijacks a conversation, and fewer times you have to repeat yourself. For Google, it's a path to personalized on-device models without requiring you to sit through a formal voice-training setup.
Google's 41st filing in voice and speech AI work we've tracked since May builds on ideas like decoding spoken account numbers and silent mid-conversation overrides.
Anyone who has a smart speaker in a kitchen or living room has experienced both failure modes: the assistant firing up because the TV said something that sounded close enough, or stubbornly ignoring a direct request three times in a row. For people who rely on voice assistants because of a disability or limited mobility, that second failure is not a minor annoyance; it is a closed door.
Google's response to this is frugal in a good way. Rather than shipping bigger models or requiring some technical setup, the system turns your natural correction ("no, I didn't mean that") into training data getting better at recognizing your voice and environment over time.
The success of the whole approach hinges on one calibration question the patent does not fully answer: how often will the system actually ask you to confirm. If those check-in prompts come too frequently, they become their own form of friction, and the fix becomes the problem.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
7 drawing sheets from US 2026/0279349 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →