Samsung Patents an AI That Compares Photos and Explains Which One Looks Better
Deciding which of two photos is sharper, better-exposed, or less noisy is surprisingly hard for a machine. Samsung's new patent describes an AI model that takes a written question about image quality, studies the photos, and responds with a reasoned answer.
What Samsung's photo-grading AI actually does
Imagine you take two shots of the same scene and ask: "Which photo has less blur?" Right now, getting a computer to answer that kind of question accurately is genuinely difficult. Most automatic quality checkers spit out a score, but they can't explain why one image is better or answer a specific follow-up question.
Samsung's patent describes an AI system that takes a plain-language question about quality, looks at the images, and produces a conversational answer. So instead of just a number, you might get: "Image 2 is sharper in the foreground but has more noise in the shadows."
The system combines three separate analyzers: one that reads your written question, one that scans the images visually, and one that measures specific quality signals like sharpness or color accuracy. All three are merged before a large language model gives its final verdict. Think of it as a photo critic who reads your brief, studies the pictures, then writes back.
… obtaining a target concatenated embedding by concatenating the target textual embedding, the target visual embedding of the at least two target images, and the target quality embedding …
Translation: The system combines text, visual data, and quality scores into a single data package.
How the model reads images, text, and quality signals together
Image quality assessment (IQA) is the science of having software judge how good a picture looks. Samsung's method adds a question-and-answer layer on top, so the system can handle prompts like "Which image has better exposure?" across at least two photos at a time.
The model has four distinct parts working in sequence:
- A text encoder reads the written quality question and converts it into a numerical representation (called an embedding) that captures its meaning.
- A visual encoder scans each image and produces a similar numerical fingerprint of what it sees.
- A quality enhancer measures specific technical indicators, things like blur level, noise, contrast, or color accuracy, and packages those as a third set of numbers.
- All three sets of numbers are stitched together ("concatenated") into one combined representation, which is then handed to a large language model (LLM), the kind of AI that powers conversational tools, to produce a written answer.
The key detail in claim 1 is that the method requires at least two images, meaning it is explicitly built for comparison tasks, not just rating a single photo. That comparative framing shapes the whole architecture.
… predicting a QA target answer for the at least two target images through a large language model (LLM) of the IQA model based on the target concatenated embedding …
Translation: An AI analyzes the combined data to generate a written explanation of which photo is better.
What this means for cameras and image editing tools
For phone cameras and photo editing apps, the ability to automatically compare shots and explain the difference in plain words could change how AI-assisted photography works. Right now, features like "Best Shot" selection pick a winner silently. A system like this could tell you why it chose one frame, or let you ask a specific question and get a direct answer.
Samsung's consistent investment in camera AI patents points to this being part of a broader effort to make its image processing pipeline more explainable, not just automatic. For consumers, the practical payoff would be a camera or editing assistant that teaches you something rather than just deciding for you.
This is the 115th Samsung filing we've tracked since May in our camera sensor push watchlist, following applications like one skipping a camera cutout and one decoding color from one signal.
Claim 1 is notable for how much it covers. It doesn't lock down a specific quality metric or a particular type of image defect. Instead, it claims the whole pipeline: text in, images in, quality signals in, LLM answer out, for any pair of images and any quality question you can phrase. That is a broad structural claim over a general-purpose comparative IQA method.
If granted as written, it could create friction for anyone building a similar question-answering wrapper around image quality analysis, whether in a camera app, a content moderation pipeline, or a photo editing tool. The concatenation step, combining three separate embedding types before the LLM, is the specific mechanical hook the claim hangs on.
The realistic question is whether prior art narrows this considerably during examination. Multimodal AI systems that merge text, visual, and quality signals already exist in the research literature. The patent will likely face pressure to tighten its claims, but even a narrowed version covering comparative multi-image IQA with an LLM output would be a meaningful piece of IP in the camera AI space.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
9 drawing sheets from US 2026/0268655 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →