New Google Patents · Filed Jan 10, 2025 · Published Jul 16, 2026 · verified — real USPTO data

Google Patent Uses Visual Search to Sharpen AI Responses to Image Queries

When you show an AI a photo and ask it a question, the AI often misses details hiding in plain sight. Google's new patent proposes a fix: run the image through a visual search engine first, then hand the AI a pre-labeled version of the photo.

Google Patent: Using Visual Search to Boost AI Image Understanding — figure from US 2026/0203363 A1
Figure from the official USPTO publication.
Publication number US 2026/0203363 A1
Applicant Google LLC
Filing date Jan 10, 2025
Publication date Jul 16, 2026
Inventors Fabio Luca Sulser, Susan Qi Xu, Vikas Bahirwani, Bhanu Prakash Reddy Guda, Lin Li, Khalid Salama, Manuel Tragut, Ágoston Weisz, Andrea Colaco
CPC classification 707/706
Grant likelihood Medium
Examiner TOUGHIRY, ARYAN D (Art Unit 2165)
Status Notice of Allowance Mailed -- Application Received in Office of Publications (Jun 3, 2026)
Document 24 claims

How Google's image-plus-search pipeline actually works

Imagine you send someone a photo of your living room and ask, "What's the name of that lamp in the corner?" A typical AI assistant might guess or confuse it with something similar. Google's patent describes a system that tackles that problem by adding an extra step before the AI even reads your question.

Here's the basic flow: your image goes into a visual search engine first (think Google Lens-style technology) which spots individual objects and draws labeled boxes around them. Then the AI gets a version of your photo with those boxes already drawn, numbered, and described in text, so it can reference specific objects by number when it answers you.

The result is that the AI doesn't have to figure out what's in the picture entirely on its own. It gets a cheat sheet, essentially, produced by a search engine that has already seen millions of similar objects. This should make the AI's answers more accurate and specific, especially for questions about particular items inside a complex scene.

How bounding boxes and reference numbers train the AI's eye

The patent describes a two-stage pipeline for answering questions about images.

Stage one: visual search engine. When a query image arrives, the system first runs it through a visual search engine (a system that matches image regions against a large index of known objects, similar to how Google Lens works). The search engine identifies objects in the scene and produces bounding shapes (boxes or outlines drawn around each detected object) along with contextual information about what those objects are.

Stage two: annotated query. The system then builds an annotated query by overlaying those bounding boxes on the original image and assigning each box a reference number. The text portion of the question is also modified to include references to those numbers, so the AI can connect "item #3" in the text to the specific outlined region in the photo.

Stage three: vision language model (VLM). A vision language model (an AI that processes both images and text together, like GPT-4o or Google's Gemini) then receives this pre-annotated package rather than the raw image. Because the hard work of locating and labeling objects has already been done, the AI can focus on reasoning about them.

The approach is essentially giving the AI a structured map of the image before it tries to answer, which reduces the chance it overlooks or misidentifies something.

What this means for Google Lens and multimodal AI search

For users, this is about getting more accurate answers when you ask questions about photos. Whether you're identifying a plant species, troubleshooting a piece of hardware from a snapshot, or comparing products by image, the system is designed to reduce the "I don't know what you're pointing at" problem that plagues current AI image tools.

For Google specifically, this matters because visual search is one of the main arenas where Google Lens, Gemini, and traditional web search are converging. A system that combines the object-recognition strengths of a search index with the reasoning ability of a large language model could meaningfully improve products like Lens or AI Overviews when they involve photos. Other companies building multimodal AI tools face the same underlying problem, so this filing signals that Google is investing in the scaffolding layer between search and AI, not just the AI itself.

Editorial take

This is a practical, un-flashy engineering fix to a real problem: AI image models are often overconfident about what they're looking at, and giving them pre-labeled maps of a scene is a sensible way to compensate. The approach of chaining a retrieval system to a generative model is well-established in text-based AI, and this is essentially Google applying that same instinct to vision. It won't get headlines, but it's the kind of infrastructure work that makes consumer products noticeably better.

Which company should we read for you?

We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.

Get one Big Tech patent every Sunday

Plain English, intelligent commentary, no hype. Free.

Source. Full patent text and figures from the official USPTO publication PDF.

Editorial commentary on a publicly published patent application. Not legal advice.