Apple · Filed May 29, 2026 · Published Sep 24, 2026 · verified — real USPTO data

Apple Patents a Way for Devices to Figure Out Exactly Which Object You're Pointing At

Point at a building and say 'that one', and your device actually knows which building you mean. Apple's new patent describes a system that combines gestures, voice, your GPS location, and a database of real-world objects to resolve exactly what you're referring to.

A device with a display showing numbered selections corresponding to real-world objects a user is pointing at. Drawing from patent filing US 2026/0287888 A1.
A device with a display showing numbered selections corresponding to real-world objects a user is pointing at.
See all 13 drawings from this filing ↓
Publication number US 2026/0287888 A1
Applicant Apple Inc.
Filing date May 29, 2026
Publication date Sep 24, 2026
Inventors Patrick S. Piemonte, Wolf Kienzle, Douglas Bowman, Shaun D. Budhram, Madhurani R. Sapre, Vyacheslav Leizerovich, Daniel De Rocha Rosario
CPC classification 345/7
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Jun 23, 2026)
Parent application is a Continuation of 18950953 (filed 2024-11-18)
Document 1 claims

What Apple's gesture-and-voice object ID system does

Imagine you're standing on a busy street corner and you point toward a row of shops and say 'navigate me there.' Today, your phone has no reliable way to figure out which shop you meant. Apple's patent takes aim at that exact frustration.

The system works by combining two clues you're already giving: a gesture (pointing, nodding, looking) or a spoken description ('the red building,' 'that café'). It then cross-references your GPS location against a database of known real-world objects nearby, narrows the list down to plausible matches, and shows them on screen for you to pick the right one without having to type or tap.

The selection itself is non-tactile, meaning you confirm your choice through a gesture or your voice rather than touching the screen. The whole loop is designed so that picking out a specific object in a crowded environment feels as natural as pointing it out to a friend.

From the filing · CLAIM 1
… obtaining a subset of real world object indicators from a database of known real world object indicators, the subset of real world object indicators being identified based on the reference to the object type and an estimated geographic location of the device …

Translation: The system narrows down nearby real world objects using your GPS location and what you mentioned.

How the system narrows down which real object you mean

The patent's first independent claim lays out a five-step process. A device receives a user request that references an object type (a shop, a landmark, a vehicle) in the surrounding environment. It then queries a database of known real-world objects, filtered by both the type you mentioned and your estimated geographic location, to return a short list of plausible matches.

That shortlist is displayed as a set of representations, likely visual overlays or a list tied to a map. The user then makes a non-tactile selection (a gesture detected by a camera or a spoken confirmation) to pick the right one. Finally, the system disambiguates the choice, meaning it resolves any remaining ambiguity and locks in the specific real-world object the user intended.

The patent describes two main input modes:

  • Visual gesture: a camera sensor detects pointing, gaze direction, or hand movement aimed at an object
  • Verbal description: a microphone captures spoken words describing the object, which the system parses to narrow candidates

The claim is written around a mobile machine context, which covers both a phone held by a person and an autonomous vehicle trying to understand a passenger's instruction. The database of known objects acts like a continuously updated map of the physical world that the system can reason against.

From the filing · THE ABSTRACT
In one example, the user may gesture to the object which is detected by a visual sensor. In another example, the user may verbally describe the object which is detected by an audio sensor.

Translation: You can either point at the object or speak out loud to describe it.

What this means for AR glasses and autonomous vehicles

This kind of interaction would be most useful in augmented reality headsets and autonomous vehicles, where touching a screen is inconvenient or impossible. If Apple's Vision Pro or a future AR product needs to let you say 'turn left at that building,' the device needs a principled way to understand which building. This patent describes exactly that pipeline.

For everyday phone users, the impact could show up in maps and navigation apps. Instead of searching by name or address, you could gesture out a window at a restaurant and ask for its hours. Apple's long bet on spatial computing makes this kind of patent a building block rather than a one-off idea, but the claim's concrete scope means it could also affect how any mobile device resolves ambiguous references to physical places.

Apple's 34th filing we've tracked since May connects to earlier applications like one that scrolls text by gaze and logging in by gaze, all part of our eye and hand controls watchlist.

Editorial take

Claim 1 is broad in a way that should get patent examiners' attention. It covers any device receiving a spoken or gestured reference to an object type, looking up nearby candidates in a database, displaying them, and resolving a non-tactile selection. That description could read on a lot of existing navigation and AR systems, which will be the main challenge during examination.

The phrase 'non-tactile selection' is the claim's sharpest edge. If granted as written, it could give Apple leverage over voice-and-gesture-based object-selection flows in competitors' AR headsets or in-car systems, not just on phones. That's a meaningful scope, not a narrow tweak.

The practical question is novelty. Voice assistants already use location to resolve ambiguous queries, and AR systems already track gaze to identify objects. Apple will need to argue that the specific combination of a candidate-shortlist display plus a non-tactile confirmation step is what makes this distinct. The claim as filed leaves room for examiners to disagree.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

13 drawing sheets from US 2026/0287888 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.