Google Patent Reveals Visual Search That Recommends Apps From Your Screen Content
Point at something on your phone screen, and Google's new patent wants the phone to figure out what it is, suggest an app, and write the message for you before you even open a keyboard.
What Google's on-screen visual search actually does
Every time you see something on your phone screen that you want to share, the path from 'I like this' to 'sent' involves a lot of tapping around. You screenshot, switch apps, paste, and maybe type a caption. Google's patent describes a way to collapse all of that into a single gesture.
The idea is an overlay that sits on top of whatever app you're already using. You draw around or tap an object on screen, and the overlay figures out what that object is. It then suggests which app you'd probably want to open next, and if that app is a messaging app, it can generate a draft message about the thing you pointed at.
You don't have to type, and you don't have to leave the app you're in. The whole flow lives inside the overlay until you decide to send something. That's the experience Google is trying to build.
… determining, by the computing system and with a machine-learned suggestion model and based on the visual search data, a suggested action comprising a create-a-text suggestion and determining, with the machine-learned suggestion model, a second application is associated with the suggested action; …
Translation: The system figures out what text to draft and which messaging app should handle it.
How the overlay spots objects and picks the right app
The patent describes a system built around an overlay application, a layer that appears on top of any other app already running on your phone. A user gesture (a tap, a circle drawn with a finger) signals which part of the screen is interesting.
From there, the system runs an on-device segmentation model (software that isolates specific objects from a larger image, similar to how a photo editor's magic lasso traces around a person) to draw a precise boundary around the object. That cropped image is then sent to a server, which performs a visual search to identify what the object is and pull up relevant information.
The system then feeds those visual search results into a machine-learned suggestion model that decides what the user might want to do next. The claim specifically names a "create-a-text suggestion," meaning the model recognizes that sharing via a messaging app is likely the right next step.
Finally, if the user picks that suggestion, a generative model (an AI that produces new text, similar to how a chatbot writes a response) drafts an actual message about the object and routes it directly to the messaging app. The entire chain, from gesture to drafted message, runs without the user switching apps manually.
The computing device can include an operating system that includes a visual search interface that obtains and processes display data associated with content currently being provided for display.
Translation: Your phone operating system looks at whatever is currently showing on your screen.
What this means for how you share things from your phone
For the average person, the payoff is time and friction. Right now, sharing something you see on a phone screen is a multi-step chore. This patent describes a way to reduce that chore to a single gesture followed by a tap to confirm. You'd notice the difference most in casual, fast moments: sending a friend a product you spotted, sharing a restaurant menu item, or flagging a piece of text without copy-pasting it.
Google's continued push into on-device AI features shows up clearly here. The segmentation step runs entirely on the phone before anything is sent to a server, which matters for speed and privacy. The question, as always with AI-drafted text, is whether the generated message actually sounds like something you'd write.
Google's 30th filing we've tracked since May on our AI agents that act for you list builds on its tap-to-identify screen buttons patent and its robot decision scoring patent to push further toward software that acts on your behalf.
The concrete payoff here is real: anyone who has fumbled through the screenshot-switch-paste cycle knows how awkward it is. A gesture-to-drafted-message flow would save genuine time in everyday situations, not edge cases.
The part that will determine whether people actually use this is the quality of the generated message. If it sounds robotic or misidentifies the object, users will delete and retype, which erases the time savings entirely. The patent doesn't say much about how the generative model is tuned for casual conversation, and that gap is where the experience will either win or lose.
On the privacy side, the segmentation runs on-device before anything leaves the phone, which is a deliberate design choice and a meaningful one. Only the cropped portion of the image gets transmitted to a server, not the full screen. That's a better setup than sending everything and sorting it out in the cloud.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
24 drawing sheets from US 2026/0259890 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →