Apple Patents a System That Shows You the Right Widget Based on What You're Looking At
Point your device at a coffee maker and get a brewing tutorial overlay. Point it at a medicine bottle and get dosage info. Apple has filed a patent for a system that figures out what you're looking at and decides which widget to show you, and when.
How Apple's object-reading widget system actually works
Imagine you're standing in your kitchen, and you hold up your phone (or glasses) toward a box of pasta. Instead of nothing happening, a little card pops up showing you a recipe, a timer, or nutritional info. That's the core idea in this Apple patent.
The device uses its camera to figure out what object is in front of you, then matches that object to a relevant widget, a small informational panel. It also tracks how you're holding the device or moving it, and that movement data determines whether the widget appears at all. If you've turned away or moved on, the widget stays hidden.
What makes this more than a simple scan-and-show system is that the device can also take in a second type of input at the same time, such as your voice or a tap, to further refine which widget it picks. Apple keeps filing on spatial computing and AR overlays, and this patent fits that pattern: getting information to appear in context, attached to the real world rather than buried in an app.
obtaining a semantic value that is associated with a physical object within a physical environment of the device; selecting a widget based on the semantic value; obtaining position data that characterizes a current pose or pose change of the device; in accordance with a determination that the position data satisfies position criteria, displaying the widget …
Translation: The system figures out what you are looking at and shows a relevant tool if you hold the device right.
How the device picks, positions, and withholds a widget
The patent describes a pipeline with a few key steps that work together.
First, the device captures image data and extracts a semantic value from it. A semantic value is essentially a label the system assigns to a physical object: "this is a coffee machine," "this is a medicine bottle," "this is a book." That label is derived from the camera feed and acts as the lookup key for choosing a widget.
Next, the system checks position data, meaning it reads the current pose (orientation and location) of the device or how that pose is changing. If the device is pointed steadily at the object, the position data satisfies the criteria and the widget is shown. If the device has moved away or is swinging past the object, the widget is suppressed. This prevents the screen from flooding with pop-ups every time the camera briefly sweeps across something.
The patent also describes a two-input version. The first input is the camera (the image modality). A second input, such as voice, touch, or another sensor, can layer additional context on top. Together, the semantic label and the secondary input data are used to select which widget to display from potentially many options.
Finally, the widget can be displayed in one of three spatial modes:
- Display-locked: anchored to a fixed spot on your screen
- Body-locked: moves with your body, like a heads-up display
- World-locked: appears to float at a fixed point in physical space relative to the object
… displaying the widget according to an object-proximity criterion (e.g., display-locked, body-locked, or world-locked) with respect to the physical object.
Translation: The interactive widget can stay anchored to your screen, your body, or a fixed spot in the real room.
What this means for AR glasses and spatial computing
For everyday users, this kind of system would make a phone or headset feel far more aware of what's actually around you. Instead of searching for information about something you see, the device reads the object and surfaces the right card automatically. That's a meaningful shift in how people interact with information in physical spaces.
The three spatial display modes are also telling. Display-locked and body-locked modes could work on a phone or tablet today. But world-locked mode, where a widget appears to hang in space at a specific real-world location, requires a device with depth sensing and spatial tracking, the kind of hardware in Apple Vision Pro. This patent describes behavior that spans both kinds of devices, suggesting Apple is thinking about a shared interaction model across its product lines.
Apple's 64th filing we've tracked since May in our AR glasses watch builds on one keeping labels on real objects and one cutting camera lag.
The basic version of this, pointing a phone camera at an object and having a relevant card appear on screen, requires no new hardware at all. The camera, the object recognition, the display logic: everything needed already exists in current devices.
The more advanced version, where information appears to float in space beside an object as you move around it, requires a device that can measure depth and track its own position in real time. That is a meaningfully different class of hardware, and the document is clear-eyed about the distinction.
The detail that determines whether either version becomes a daily habit is the rule about suppressing cards unless the device is actually pointed at the target object. Without that filter, lifting your phone triggers a flood of information you did not ask for, and the feature becomes something people turn off in the first week.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
20 drawing sheets from US 2026/0288301 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in