Apple · Filed Mar 27, 2026 · Published Aug 13, 2026 · verified — real USPTO data

Apple Patents a Way to Figure Out What You Mean by Watching Where You Look

When you say 'turn that off,' your smart speaker has no idea what 'that' is. Apple's new patent proposes using eye-tracking and environmental scanning to figure it out, turning a vague spoken command into a precise action.

A device tracking a user's gaze direction to identify and select specific objects in the surrounding area. Drawing from patent filing US 2026/0237390 A1.
A device tracking a user's gaze direction to identify and select specific objects in the surrounding area.
See all 4 drawings from this filing ↓
Publication number US 2026/0237390 A1
Applicant Apple Inc.
Filing date Mar 27, 2026
Publication date Aug 13, 2026
Inventors Kenneth M. KARAKOTSIOS, James BYUN, Pulah J. SHAH
CPC classification 704/275
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (May 5, 2026)
Parent application is a Continuation of 18227884 (filed 2023-07-28)
Document 15 claims

How Apple's gaze-plus-voice system resolves what you mean

Every time you tell a voice assistant to do something, it has to guess what you're talking about from words alone. Say 'make it brighter' in a room full of smart lights and connected devices, and the assistant is essentially guessing.

Apple's patent describes a system where a camera or sensor first scans your surroundings to build a list of nearby objects. Then, a second sensor (most likely an eye-tracker in a headset or glasses) notices where you're looking. When you speak a command, the system already knows which object you're probably talking about, because it followed your gaze. The result is that 'turn that off' finally means something specific.

This approach links three things together: a physical scan of your environment, your eye movements, and your spoken words. Together, they let a device resolve the kind of everyday ambiguity that trips up even the best voice assistants today.

From the filing · CLAIM 1
… analyzing the scan of the proximate area to identify objects in the proximate area; after analyzing the scan of the proximate area, detecting, with a second sensor of a second external electronic device different from the first external electronic device, a first user input corresponding to at least one of the identified objects …

Translation: The system scans your surroundings to identify items, then uses a second sensor to see which one you are interacting with.

How the two-sensor pipeline connects gaze to commands

The patent describes a two-phase process that runs before you even finish speaking.

Phase one involves a first external device (think a spatial camera or room scanner) sweeping the nearby area and identifying every recognizable object it can find. This builds a kind of inventory of your environment.

Phase two is triggered by your gaze. A second external device, such as a head-mounted display with eye-tracking, detects where you're looking. That gaze location acts as a filter: the system narrows its focus to whichever object or objects fall in your line of sight, then requests detailed information about just those items from an external source.

When you then speak a command, the device already has:

  • A map of nearby objects
  • A record of which one you were looking at
  • Contextual information about that specific object

The spoken command is then interpreted in light of all that context, allowing the device to perform the right action even when the words themselves are ambiguous. The patent is careful to separate the scanning sensor from the gaze sensor, suggesting Apple envisions this working across multiple hardware devices acting in coordination.

From the filing · THE ABSTRACT
Input regarding the user, such as a user’s gaze location, may trigger a second analysis of a subset of the identified objects in the environment. The results of these analyses may then be used to resolve an ambiguity in linguistic user input from the user.

Translation: The device tracks where you are looking to figure out which object you are talking about when your commands are unclear.

What this means for Siri and spatial computing

For Siri, this could be the architectural fix that makes spatial voice control actually feel natural. The biggest frustration with voice assistants in multi-device homes or mixed-reality environments is that language is inherently imprecise. A system that fuses gaze, environment, and speech data sidesteps that problem at the source rather than trying to patch it with smarter language models alone.

Apple's Vision Pro already ships with both eye-tracking and spatial awareness hardware, which puts this patent closer to real infrastructure than a blue-sky idea. Readers interested in where this fits among interesting tech patents covering AR and spatial computing will notice Apple has been building each layer of this stack separately, and this filing is the first to stitch gaze explicitly into command resolution.

Editorial take

The hardware prerequisites for this patent already exist inside Vision Pro: an eye-tracking system, outward-facing cameras, and a processing layer that can fuse sensor data. That means the path from filing to shippable feature is shorter here than for most spatial-computing patents, which typically depend on sensors that haven't reached consumer devices yet. The main open question is whether Apple can run this pipeline fast enough to feel instantaneous rather than like the device is thinking. If latency is solved, this could land as a visionOS update rather than requiring entirely new hardware.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

4 drawing sheets from US 2026/0237390 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.