Nvidia · Filed Mar 6, 2026 · Published Sep 10, 2026 · verified — real USPTO data

Nvidia Patents a System That Reads Your Gaze and Offers Device Controls Instantly

Nvidia has filed a patent for a system that watches where your eyes land on a scene, figures out what object you're looking at, and immediately offers a menu of things you can do with it. No typing, no navigating settings, no fumbling for a remote.

A system processes image data, attention signals, and environmental data to present ranked task identifiers and receive user commands. Drawing from patent filing US 2026/0267458 A1.
A system processes image data, attention signals, and environmental data to present ranked task identifiers and receive user commands.
See all 28 drawings from this filing ↓
Publication number US 2026/0267458 A1
Applicant NVIDIA Corporation
Filing date Mar 6, 2026
Publication date Sep 10, 2026
Inventors Julien Francois Jomier, Mahdi Azizian, Nigel Scott Nelson
CPC classification 715/708
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Apr 21, 2026)
Parent application Claims priority from a provisional application 63768603 (filed 2025-03-07)
Document 20 claims

What Nvidia's eye-gaze task suggestion actually does

Imagine glancing at your thermostat on the wall and having a small menu pop up in your field of view offering "Turn up," "Turn down," or "Set schedule" the moment your eyes land on it. That's the core idea here: your gaze tells the system what you care about, and the system does the thinking about what you might want to do next.

Nvidia's patent describes a setup where a camera or sensor tracks exactly where you're looking within a visual scene. An AI model then figures out which physical object is at that spot and generates a short list of relevant actions tied specifically to that device or item. Pick one from the list, and the system sends the command directly to that device.

The appeal is that you skip several steps that normally sit between noticing something and doing something about it. You don't have to open an app, find the right device in a list, or remember what a physical button does. Your attention itself becomes the input.

From the filing · CLAIM 1
receive image data representing a visual scene and spatial telemetry data comprising a gaze coordinate indicating a location within the visual scene …

Translation: It tracks where you are looking inside a live video feed of your surroundings.

How gaze data and AI combine to trigger device commands

The system works in three stages that happen in rapid sequence.

First, it captures image data of whatever visual scene is in front of you, along with spatial telemetry, which is a precise coordinate pinpointing where your eyes are focused within that scene. That coordinate is called a gaze coordinate, essentially a map pin dropped on whatever you're looking at.

Second, both inputs are fed into a vision-language model (a type of AI trained to understand images and describe them in natural language). The model cross-references the gaze coordinate with the image to identify the specific object at that location. From there it generates task identifiers, short labels representing actions that make sense for that object, such as "dim," "lock," "play," or "increase temperature."

Third, when you select one of those task identifiers, the system sends a command through a device control interface tied to that object. That interface is essentially a software bridge between the AI layer and the physical or networked device being controlled.

The patent covers both the core logic and the adaptive presentation side, meaning the system can adjust how and where those task options appear based on context.

From the filing · THE ABSTRACT
… process the image data and the attention signal using a vision-language model to identify an object corresponding to the location, and generate task identifiers representing tasks associated with the object …

Translation: An AI figures out what you are staring at and lists the actions you can take on it.

What this means for hands-free and AR device control

For anyone using augmented reality glasses, smart home dashboards, or surgical or industrial control systems, this kind of interaction could remove a lot of friction. Instead of voice commands (which require remembering the right phrase) or touch screens (which require finding the right app), your visual focus does the navigation for you. That's a meaningful shift for situations where your hands are busy or unavailable.

Nvidia's steady investment in vision-language and robotics filings suggests this isn't a one-off idea. The gaze-driven control concept maps naturally onto robot teleoperation, where an operator might look at a robot's camera feed and want to act on whatever object the robot is near, as well as onto consumer AR, which is still searching for its killer interaction model.

This is the fourth Nvidia filing we've tracked since July in the AR glasses race, joining earlier applications on eye-tracking without real faces and putting remote users in your room.

Editorial take

Controlling a smart device today means remembering what it's called in an app, finding it, and then picking what you want it to do. This patent replaces all of that with simply looking at the thing.

The experience only holds up if the system correctly identifies what you're looking at. Wrong object, wrong menu, and the whole shortcut evaporates.

The people who benefit most immediately are those who cannot spare a hand: surgeons, factory workers, people operating remote equipment, where fumbling for a phone creates real risk. For everyone else, it is a feature waiting on hardware most of us do not own yet.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

28 drawing sheets from US 2026/0267458 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.