Google Patents a Voice Assistant That Can Navigate to Places You're Not Even Looking At
Most voice navigation only works with the spot you're already looking at on the map. Google's new patent describes a system that can find and navigate to a place that's completely off your screen, just from a vague phrase like 'that area' or 'over there.'
What Google's off-screen voice navigation actually does
Imagine you're looking at a map of your neighborhood and you say to your phone, "Navigate to that park over there" while vaguely gesturing toward a part of the city not even on screen. Right now, your phone would probably have no idea what you mean. That's the gap this patent is trying to close.
Google's system listens for words like "there," "that place," or "over there" (called distal words, meaning words that point to something distant or out of view) and uses them to search for matching destinations outside the map area you're currently looking at. It then pulls up the right spot and shifts the map to show it to you.
The goal is a more natural way to interact with navigation, especially when you have a rough idea of where you want to go but haven't scrolled the map there yet. Instead of typing or zooming around, you just say something vague and let the assistant figure out the region you mean.
… determining that the identified referential word is classified as a distal word; in response to determining that the referential word is classified as the distal word, identifying a subset of point locations corresponding to a second geographic region that is currently outside the viewport of the navigation application …
Translation: The system figures out when you are asking about a place that is currently hidden off screen.
How the system links 'over there' to a point outside your view
The patent describes a data processing system that sits between a voice assistant and a navigation app, coordinating information between the two.
Here's how the pieces fit together:
- The navigation app shares a list of point locations (named places, landmarks, addresses) along with their coordinates and identifiers, covering not just what's visible on screen but a broader geographic database.
- When you speak, the system parses the audio to find two things: a request ("navigate to," "take me to") and a referential word ("there," "that area," "over there").
- It then classifies that referential word as a distal word, meaning it refers to something outside the current map view. This is the key branching point: if the word pointed to something visible on screen, a different path would apply.
- Using that classification, it searches point locations in a second geographic region that matches where the user seems to be pointing or referring to, identifies the best match, and sends the navigation app a structured instruction to display and route to that spot.
The system essentially translates spatial language (words that gesture at a direction or vague region) into map coordinates, then hands those coordinates off to the navigation app as a formal action.
The data processing system can parse an input audio signal to identify a request and a referential word.
Translation: Your voice command is analyzed to figure out what you want and which location words you used.
What this means for hands-free map use while driving
For drivers and cyclists, this is about reducing friction. Fumbling with a map to scroll to a region you vaguely remember, just so you can tap a destination, is annoying and dangerous at the wheel. A system that interprets your rough verbal gesture and finds the right place automatically removes several steps from that process.
The design also reflects a broader challenge in voice interfaces: most of them treat spoken language too literally, requiring you to say exact names or addresses. Google keeps filing on natural-language interaction with device interfaces This patent bets that making assistants understand approximate spatial references, the way a human passenger would, is worth the engineering complexity of linking two separate apps in real time.
Google's 699th filing we've tracked since May in our Google coverage continues a thread of AI listening ideas, joining a self-correcting voice system and a catch-up assistant session.
The system bets everything on correctly interpreting a vague word like "there" or "that area" before committing to a route. When it guesses wrong, the consequence is not a mildly annoying screen, it is a wrong turn in the real world.
The document says nothing about what happens when several plausible destinations exist nearby, or whether the system signals uncertainty before locking anything in. Those are real gaps, and they represent the actual cost of this approach: the experience only works as well as the system's ability to read your mind on the first attempt.
Still, the underlying trade reads as worth making. Voice navigation today forces people to speak in ways nobody naturally talks, demanding precise names instead of loose, conversational language. Teaching a system to meet people where they are is a sensible direction, as long as the failure cases are handled with more care than this document suggests.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
6 drawing sheets from US 2026/0279347 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in