New Google Patents · Filed Jun 1, 2026 · Published Oct 1, 2026 · verified — real USPTO data

Google Patents a System That Learns When Its Voice Assistant's Follow-Up Questions Aren't Working

When you ask a voice assistant something ambiguous and it asks a confusing follow-up question, the whole conversation falls apart. Google is patenting a system that tracks when those follow-up questions fail and automatically upgrades them to something more useful.

An assistant device communicates with a cloud-based automated assistant to process user requests and manage data. Drawing from patent filing US 2026/0301740 A1.
An assistant device communicates with a cloud-based automated assistant to process user requests and manage data.
See all 4 drawings from this filing ↓
Publication number US 2026/0301740 A1
Applicant GOOGLE LLC
Filing date Jun 1, 2026
Publication date Oct 1, 2026
Inventors Matthew Sharifi, Victor Carbune
CPC classification 704/257
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Jun 27, 2026)
Parent application is a Continuation of 18244762 (filed 2023-09-11)
Document 18 claims

How Google's assistant decides to ask better questions

You're in the kitchen and you tell your Google Assistant to "play something relaxing" and it can't tell if you want music or a podcast. So it asks a clarifying question. That question is either helpful or baffling, depending on how it's delivered.

Right now, voice assistants mostly just ask again in plain speech: "Did you mean music or a podcast?" That works sometimes. But Google is filing a patent for a system that keeps score on whether plain-speech clarifications are actually solving the problem. If they're not, the assistant automatically switches to a richer format: think visual options on a screen, or audio examples, anything beyond just more words.

The key idea is that the assistant isn't guessing randomly. It's analyzing its own track record. When plain questions have a history of failing in similar situations, the system flags that and chooses a more helpful format instead.

From the filing · CLAIM 1
determining whether, in soliciting further user input for disambiguating between the first particular action and the second particular action, to provide a natural language (NL) only clarification prompt that is restricted to rendering NL, or to instead provide an enhanced clarification prompt that renders NL and renders output that is in addition to NL …

Translation: The system decides whether to ask a clarifying question using only speech or to also use a screen or other outputs.

How the failure metric triggers a richer clarification prompt

The patent describes a method for handling ambiguous voice commands in two stages: detecting the ambiguity, then deciding how to ask for clarification.

When a user says something that could mean two mutually exclusive things, the assistant must ask for more input. But instead of always defaulting to a plain-language question ("Did you mean X or Y?"), the system evaluates a failure metric: a score built from historical data about how often that type of plain-language clarification failed to resolve the ambiguity in past interactions.

  • If the failure metric is below a threshold, the assistant sticks with a plain-language question (cheaper to render, less disruptive).
  • If the failure metric meets or exceeds the threshold, the assistant switches to an enhanced clarification prompt that layers in additional output beyond speech, such as a visual selection menu or audio clips.

The system also factors in prior determinations: if the assistant has already decided to use enhanced prompts in a recent session, that history can influence the current choice. The underlying failure metric is continuously shaped by aggregate interaction data, not just a single user's history.

From the filing · THE ABSTRACT
… determine, based on processing the recognition, that the spoken utterance is ambiguous (i.e., is interpretable as requesting performance of a first particular action exclusively and is also interpretable a second particular action exclusively).

Translation: The assistant realizes it cannot understand a command because it points to two different actions at the same time.

What this means for people who argue with their smart speakers

For anyone who uses a smart speaker or a phone assistant daily, this is a fix for one of the most frustrating moments in that relationship: when the device asks you a question and you get confused. Plain verbal clarifications often fail because they add more words to an already murky situation.

the pattern in Google's assistant-interaction filings suggests the company is increasingly interested in making the assistant's failure modes smarter rather than just making the assistant less likely to fail in the first place. Whether this shows up in Google Home devices, Android phones, or in-car assistants, the practical result would be fewer dead-end exchanges where you end up just repeating yourself louder.

Google's 48th filing in voice and speech AI we've tracked since May follows an always-same-speed audio generator and a turn-detection voice system.

Editorial take

The core trade here is between simplicity and accuracy. A plain-language follow-up is lightweight: no screen needed, no extra latency, works on a voice-only device. An enhanced prompt is heavier: it requires a display or richer audio capability, and it interrupts the conversational flow more visibly. Google's system tries to make that call adaptively, only paying the cost of a richer prompt when history says the cheaper option doesn't work.

The risk is in how that failure metric gets defined and measured. If the system marks a clarification as "failed" only when the user explicitly corrects the assistant, it will miss quieter failures: the person who just gives up, or accepts the wrong answer. Under-counting failures means the threshold gets hit less often, and the system stays on plain-language prompts longer than it should.

That said, the design logic is sound for the devices where it's most needed. A smart display already has the hardware to show options; the problem was never capability, it was knowing when to use it. Teaching the assistant to make that call based on its own track record is a more honest approach than just showing visual menus all the time and calling it an improvement.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

4 drawing sheets from US 2026/0301740 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.