Samsung · Filed Nov 13, 2025 · Published Jul 16, 2026 · verified — real USPTO data

Samsung Patents an Image Search That Translates Your Photos Into Words to Find Better Matches

When you search by photo on your phone today, you're limited to what visual similarity can find. Samsung's new patent adds a second layer: it turns your image into a text description and searches that way too, then combines both results.

Samsung Patent: AI Image Search Using Text and Photo Together — figure from US 2026/0203342 A1
Figure from the official USPTO publication.
Publication number US 2026/0203342 A1
Applicant Samsung Electronics Co., Ltd.
Filing date Nov 13, 2025
Publication date Jul 16, 2026
Inventors Ilwi YUN, Seongeun KIM, Seungin PARK, Dongwook LEE, Wonjun CHOI, Sunjun HWANG
CPC classification 707/741
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Dec 16, 2025)
Document 20 claims

How Samsung's photo-plus-text search actually works

Imagine you snap a photo of a lamp you love at a friend's house and want to find something similar online. Today, most image search tools compare your photo to other photos visually, matching colors and shapes. That works sometimes, but it misses a lot.

Samsung's patented approach adds a twist: it takes your photo and automatically writes a text description of it, something like "minimalist white ceramic table lamp with a linen shade." Then it runs two searches at once, one using your original photo and one using that generated description. Both searches pull back a pool of candidate images.

A second AI model then looks at all of those candidates together, alongside your original photo and the generated text, and scores them. The images with the highest combined score are the ones you actually see. The idea is that combining visual and text signals catches things either search alone would miss.

Inside Samsung's two-model retrieval and scoring pipeline

The system runs in two distinct stages.

In the first stage, a trained model (the patent calls it the "first model") takes your image query and generates a text description of it automatically. That text query is then used to search a database for matching text documents, while the original image is used in a parallel visual search. The result is two separate pools of candidates: images that look visually similar and images linked to text that matches the description.

In the second stage, the patent introduces a multimodal scoring model (a model that understands both images and text simultaneously). This second model receives your original photo, the generated text query, and the full candidate image pool. It assigns each candidate a multimodal score, essentially a single number reflecting how well each image matches both the visual and textual intent of your search.

  • First model: image in, descriptive text out
  • Database lookup: parallel image search and text search
  • Second model: reranks the merged candidate pool using both signals
  • Output: the top-scoring images presented to the user

The key insight is that neither visual search nor text search alone is consistently reliable. Combining their outputs through a learned scoring step aims to surface more relevant results than either method would on its own.

What this means for photo search on Galaxy devices

For anyone who has used Google Lens or Samsung's own Bixby Vision and been frustrated when visually similar but contextually wrong images appear, this approach targets exactly that problem. By injecting a text understanding layer into what looks like a simple photo search, the system can catch nuances that pure image similarity misses, like the difference between "modern ceramic lamp" and "antique brass chandelier" when both have similar visual profiles.

For Samsung, this kind of search capability fits naturally into Galaxy phone features, particularly in the gallery app or any shopping-linked visual search tool. It is also relevant to any service where users upload an image and expect the system to understand what they are looking for, not just what it looks like.

Editorial take

This is a solid, practical improvement to a real problem with image search, not a flashy AI demo. The two-stage approach (generate text from image, search both, then rerank with a combined model) is the kind of engineering refinement that makes a feature actually useful. Whether Samsung ships it in a Galaxy update or uses it behind the scenes in a shopping or gallery feature, it is worth watching.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.