New Google Patents · Filed Mar 16, 2026 · Published Jul 23, 2026 · verified — real USPTO data

New Google Patent Trains Image Search Models on Real User Clicks

Instead of teaching an AI what 'similar images' means by hand, Google wants to let millions of real user clicks do the teaching, turning everyday search behavior into a training signal for its visual search engine.

Google Patent: Training AI to Find Visually Similar Images — figure from US 2026/0211940 A1
Figure from the official USPTO publication.
Publication number US 2026/0211940 A1
Applicant Google LLC
Filing date Mar 16, 2026
Publication date Jul 23, 2026
Inventors Zhen Li, Yi-Ting Chen, Yaxi Gao, Da-Cheng Juan, Aleksei Timofeev, Chun-Ta Lu, Futang Peng, Sujith Ravi, Andrew Tomkins, Thomas J. Duerig
CPC classification 706/12
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Apr 17, 2026)
Parent application is a Continuation of 18741082 (filed 2024-06-12)
Document 20 claims

How Google's click-based image search training works

Imagine you're searching Google Images for a red velvet cake recipe and you keep clicking two different photos that look nearly identical. Google notices that pattern. When lots of people click the same pair of images during a single session, that's a signal that those two photos are meaningfully similar, even if the AI didn't already know that.

This patent describes a system where Google uses those click patterns to train its image-recognition AI. Instead of programmers writing rules about what makes two pictures 'alike,' the system learns from the crowd: if users consistently treat two images as interchangeable, the model treats them as close matches too.

The result is an image search engine that gets better at understanding what you actually want when you search by photo, not just what the pixels technically have in common. Think of it as crowd-sourced taste teaching the algorithm.

How co-click rates shape the image embedding model

The patent covers a method for training what's called an image embedding model (a system that converts a photo into a list of numbers representing its visual meaning, so that similar photos end up with similar numbers).

The training data isn't hand-labeled by humans or scraped from image metadata. Instead, Google uses two specific behavioral signals from its search logs:

  • Co-click rate: how often two images are clicked together in the same search session, suggesting users see them as equivalent choices
  • Similar-image click rate: how often users click a suggested 'similar image' link connecting two photos, a direct vote that the images feel alike

At inference time (when the system is actually running), a query image (the photo you search with) is converted into an embedding by the trained model. That embedding is then compared against a database of other image embeddings using similarity scores (basically, mathematical distance between the number-lists). The closest matches become your search results.

The key innovation is using implicit user behavior, rather than explicit labels, as the ground truth for what 'similar' means.

What this means for Google Lens and visual search

Google Lens and Google Images already let you search by photo, but the hard part is deciding what 'similar' really means to a human, not just a pixel-counter. A photo of a beige sofa and a photo of an ivory sofa might look different to a color histogram but identical to a furniture shopper. Click-pattern training captures that human judgment at scale, without Google needing to pay annotators to label billions of image pairs.

For you as a user, a well-trained model like this means fewer frustrating results when you snap a photo of a product and ask Google to find it. For Google, it's a way to keep improving visual search quality using data it already has in enormous quantities.

Editorial take

This is solidly useful infrastructure work, not a headline-grabbing concept. Using implicit user behavior as a training signal is a well-established idea in recommendation systems, and applying it to image embeddings is a natural extension. What makes it worth noting is the scale: Google's click logs are enormous, and that data advantage is hard for competitors to replicate.

Which company should we read for you?

We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.

Get one Big Tech patent every Sunday

Plain English, intelligent commentary, no hype. Free.

Source. Full patent text and figures from the official USPTO publication PDF.

Editorial commentary on a publicly published patent application. Not legal advice.