Samsung · Filed Nov 13, 2025 · Published Jul 16, 2026 · verified — real USPTO data

Samsung Patents an Image Search That Translates Your Photos Into Words to Find Better Matches

When you search by photo on your phone today, you're limited to what visual similarity can find. Samsung's new patent adds a second layer: it turns your image into a text description and searches that way too, then combines both results.

Samsung Patent: AI Image Search Using Text and Photo Together — figure from US 2026/0203342 A1
Figure from the official USPTO publication.
Publication number US 2026/0203342 A1
Applicant Samsung Electronics Co., Ltd.
Filing date Nov 13, 2025
Publication date Jul 16, 2026
Inventors Ilwi YUN, Seongeun KIM, Seungin PARK, Dongwook LEE, Wonjun CHOI, Sunjun HWANG
CPC classification 707/741
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Dec 16, 2025)
Document 20 claims

How Samsung's photo-plus-text search actually works

Imagine you snap a photo of a lamp you love at a friend's house and want to find something similar online. Today, most image search tools compare your photo to other photos visually, matching colors and shapes. That works sometimes, but it misses a lot.

Samsung's patented approach adds a twist: it takes your photo and automatically writes a text description of it, something like "minimalist white ceramic table lamp with a linen shade." Then it runs two searches at once, one using your original photo and one using that generated description. Both searches pull back a pool of candidate images.

A second AI model then looks at all of those candidates together, alongside your original photo and the generated text, and scores them. The images with the highest combined score are the ones you actually see. The idea is that combining visual and text signals catches things either search alone would miss.

Inside Samsung's two-model retrieval and scoring pipeline

The system runs in two distinct stages.

In the first stage, a trained model (the patent calls it the "first model") takes your image query and generates a text description of it automatically. That text query is then used to search a database for matching text documents, while the original image is used in a parallel visual search. The result is two separate pools of candidates: images that look visually similar and images linked to text that matches the description.

In the second stage, the patent introduces a multimodal scoring model (a model that understands both images and text simultaneously). This second model receives your original photo, the generated text query, and the full candidate image pool. It assigns each candidate a multimodal score, essentially a single number reflecting how well each image matches both the visual and textual intent of your search.

  • First model: image in, descriptive text out
  • Database lookup: parallel image search and text search
  • Second model: reranks the merged candidate pool using both signals
  • Output: the top-scoring images presented to the user

The key insight is that neither visual search nor text search alone is consistently reliable. Combining their outputs through a learned scoring step aims to surface more relevant results than either method would on its own.

What this means for photo search on Galaxy devices

For anyone who has used Google Lens or Samsung's own Bixby Vision and been frustrated when visually similar but contextually wrong images appear, this approach targets exactly that problem. By injecting a text understanding layer into what looks like a simple photo search, the system can catch nuances that pure image similarity misses, like the difference between "modern ceramic lamp" and "antique brass chandelier" when both have similar visual profiles.

For Samsung, this kind of search capability fits naturally into Galaxy phone features, particularly in the gallery app or any shopping-linked visual search tool. It is also relevant to any service where users upload an image and expect the system to understand what they are looking for, not just what it looks like.

Editorial take

This is a solid, practical improvement to a real problem with image search, not a flashy AI demo. The two-stage approach (generate text from image, search both, then rerank with a combined model) is the kind of engineering refinement that makes a feature actually useful. Whether Samsung ships it in a Galaxy update or uses it behind the scenes in a shopping or gallery feature, it is worth watching.

Which company should we read for you?

We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.

Get one Big Tech patent every Sunday

Plain English, intelligent commentary, no hype. Free.

Source. Full patent text and figures from the official USPTO publication PDF.

Editorial commentary on a publicly published patent application. Not legal advice.