Qualcomm · Filed Feb 14, 2025 · Published Aug 20, 2026 · verified — real USPTO data

Qualcomm Patents AI That Cuts Out Objects in Video Using Text and Image Cues Together

Qualcomm has filed a patent for an AI system that can isolate objects in video frames by accepting two different kinds of instructions at the same time, such as a typed description and a tapped region, rather than forcing you to pick just one.

Visual comparison of object isolation in an image using text, pointing cues, or both combined. Drawing from patent filing US 2026/0245334 A1.
Visual comparison of object isolation in an image using text, pointing cues, or both combined.
See all 8 drawings from this filing ↓
Publication number US 2026/0245334 A1
Applicant QUALCOMM Incorporated
Filing date Feb 14, 2025
Publication date Aug 20, 2026
Inventors Taotao JING, Shuai ZHANG, Eyasu Zemene MEQUANINT, Yingyong QI
CPC classification 382/156
Grant likelihood Medium
Examiner KOROMA, SORIE IBRAHIM (Art Unit 2662)
Status Docketed New Case - Ready for Examination (Mar 14, 2025)
Document 20 claims

What Qualcomm's mixed-input video cutout system actually does

A security camera stares at an empty hallway all night. Eventually a person walks in, and you want the system to highlight just that person, not the whole frame. Today, most AI tools make you choose how to tell them what to find: either you draw a box, or you type a label, but rarely both at once.

Qualcomm's patent describes a system that accepts two different kinds of input at the same time. You could, for example, tap a spot on the screen and type "person in a red jacket," and the AI would use both clues together to paint a precise outline around the right target. The result is a highlighted mask, a colored overlay, drawn right over the video frame.

The whole process is designed to run on the kind of chip Qualcomm builds for phones and cameras, not on a distant server. That matters because it keeps your video local and makes the response nearly instant.

From the filing · CLAIM 1
… generate, by a neural network trained to perform mask segmentation based on a plurality of prompts, one or more masks for one or more target regions, in the frame, determined based on: the first prompt; the second prompt; and the plurality of single-scale features …

Translation: An AI creates precise cutouts using multiple types of instructions and video features combined.

How the neural network combines two prompt types into one mask

The patent describes an apparatus, most likely a chip or system-on-chip, whose processing system runs a neural network trained specifically for mask segmentation (drawing pixel-precise outlines around objects in an image or video frame).

What makes this different from standard segmentation is the input side. The system takes two "prompts" from two different modalities (types of input), such as:

  • A text description, for example "the dog near the door"
  • A visual cue, for example a tapped point, a drawn box, or a reference image crop

Both prompts are fed into the neural network simultaneously. The network also pulls what the patent calls single-scale features from the frame, meaning it extracts a single, fixed-resolution representation of the image rather than processing it at multiple zoom levels. Single-scale extraction is computationally lighter, which is important for running on mobile hardware.

The network combines all three inputs (prompt one, prompt two, and the frame features) and outputs one or more masks, colored overlays that mark the target region. The final frame, with the mask applied, is then sent to the display. The design is explicitly built to handle multiple prompts, not just one, making the segmentation more precise when the user's intent is ambiguous from a single signal alone.

From the filing · THE ABSTRACT
… obtaining a first prompt associated with a first modality; obtaining a second prompt associated with a second modality that is different than the first modality …

Translation: The system accepts two different input types like text and images to guide the editing process.

What this means for AI cameras and on-device processing

Qualcomm makes the processors that go inside Android phones, security cameras, AR headsets, and automotive systems. A segmentation method designed to run efficiently on those chips, without sending video to the cloud, has practical applications across all of those product categories. For you as a user, it could mean your phone's camera app highlights exactly the subject you want edited, or a smart-home camera flags the specific person you described, faster and without a network connection.

The ability to combine a typed description with a visual pointer is the more interesting part of the claim. Segmentation tools that accept only one type of input can misfire when a scene is crowded or ambiguous. Using two modalities together narrows the target more reliably. Qualcomm's chip division has been filing steadily in on-device AI, and this patent joins a broad stream of interesting tech patents focused on running complex vision tasks directly on mobile and edge processors rather than in the cloud.

Qualcomm's sixth filing we've tracked in our AI photo editing race watchlist since July builds on one rendering parts at different resolutions and one exposing objects separately.

Editorial take

The patent claim covers any device that takes two different types of input (say, a text description plus an image) and draws an outline around the object in question. That is a wide net. It could cover everything from a phone camera to a doorbell camera, as long as the device follows that basic recipe.

The one thing that narrows it down is a rule about keeping the processing simple and light, rather than running the heavy, layered analysis that cloud servers typically use. That detail pushes the claim toward products that do their work right on the device itself, and likely leaves out server-based systems that rely on those heavier methods.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

8 drawing sheets from US 2026/0245334 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.