Qualcomm Patents a Way to Teach On-Device AI Cameras to Spot Objects They've Never Seen Before
Most AI cameras can only recognize what they were trained to see before they shipped. Qualcomm is working on a way to teach them new objects on the fly, without sending data back to a server.
What Qualcomm's new visual-labeling AI actually does
A security camera stares at an empty hallway all night. It can spot a person walking by, but if your warehouse gets a new type of forklift or a new style of protective helmet, the camera's AI has no idea what it's looking at, it just sees a blob. That gap is what this patent is designed to close.
Qualcomm's approach lets someone type a label, like "hard hat" or "new forklift model," and pair it with the specific pixels in a video frame that show that object. The AI then updates its understanding so it can recognize that object on its own going forward.
The whole process is designed to happen on the device itself, meaning on the chip inside the camera or robot, rather than in a cloud data center. That matters for situations where you need a fast response, have a weak internet connection, or are handling footage that shouldn't leave the building.
… train, a text encoder of the neural network, to associate a text embedding of the first text label with a mask embedding of the first mask.
Translation: The system learns to link specific words to the visual shapes of objects so it can identify them later.
How the text encoder links words to pixel-level masks
The patent describes a training method for a type of AI called a mask-based segmentation model, an AI that doesn't just draw a box around an object but traces its exact outline in a frame, down to the pixel level.
When the AI analyzes a frame, it produces a mask (a precise pixel-by-pixel silhouette of a region) and a mask embedding (a compressed numerical fingerprint of what that region looks like). Separately, a text encoder converts a human-typed label like "cracked pipe" into its own numerical fingerprint called a text embedding.
The patent's core idea is to train the text encoder so these two fingerprints line up: the numbers that represent the word "cracked pipe" should land close to the numbers that represent what a cracked pipe looks like in pixels. Once they're aligned, the model can match new text queries to objects in future frames, even for categories it never saw during its original training.
- A user (or another system) provides a text label for a detected region.
- The neural network links the text embedding to the mask embedding through a training step.
- Future frames can then be segmented using that new vocabulary word as a query.
The present disclosure provide techniques for mask-based frame segmentation. A method may include obtaining, based on input, a first text label related to a first mask, generated by a neural network, for a first target region depicted in a first frame; …
Translation: This technology uses text descriptions to help the camera isolate and identify specific objects within a video frame.
What this means for cameras that run AI on-chip
For anyone deploying cameras or sensors in environments that change, a factory floor that adds new equipment, a construction site with evolving safety gear, a retail space that cycles through new product packaging, the usual workaround is to ship the device back to a vendor or wait for a model update. This patent describes a path where a technician or an automated system could just label an example on the spot and have the device learn from it immediately.
Qualcomm makes the chips that power a large share of the world's cameras, phones, and edge-computing devices, so a technique like this, if it makes it to production, could land in a wide range of hardware. The AI-on-chip space is one of the more active areas for new Big Tech patents, with camera vendors, chipmakers, and robotics companies all racing to make devices that adapt without a round-trip to the cloud.
That makes this Qualcomm's eighth filing since July on our on-device AI privacy watchlist, following work on AI camera power saving and keeping location data fresh.
If you run a factory floor or a security operation, you've hit the wall where an AI camera confidently tracks everything it was trained on, then goes completely blank when something new rolls in. This patent targets that exact moment, and it fixes it at the labeling step, which is where the real-world time drain lives.
The footage never leaves the device. That matters because it means no waiting on a server, no sending sensitive video across a network, and faster response when something unfamiliar needs to be identified and acted on.
Whether this reaches an actual product depends on decisions Qualcomm hasn't announced, and patents don't predict those reliably. But the failure it prevents is one that operators notice immediately, and that's a reasonable place to start.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
7 drawing sheets from US 2026/0253367 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →