Samsung Files Patent for AI That Maps Who and What Is in Your Photos
Samsung is training an AI to look at a photo and understand not just what's in it, but how everything in the scene relates to everything else. That's a harder problem than it sounds, and the answer could reshape how your phone's camera roll actually works.
What Samsung's photo-scene-reading AI actually does
Imagine you take a photo at a birthday party: your friend is blowing out candles on a cake. Your phone can probably tag "person" and "cake." But can it understand that the person is interacting with the cake, or that a second person in the background is watching? That kind of scene-level understanding is what Samsung's patent is after.
The idea is to train an AI model using photos, their text captions, and the built-in knowledge of a large language model (the same kind of AI that powers chatbots). By combining visual examples with language-based common sense, the system learns to label not just what things are in an image, but how they relate to each other.
For you as a user, this could eventually mean a photo search that finds "photos where I'm cooking" or a camera that automatically understands context without you typing a single word.
… training, using the at least one processing device, a machine learning model to determine entity categories and relations between entities in input images based on the training dataset, wherein the entities include subjects and objects captured in the input images.
Translation: The system learns to identify people and items in your photos and understands how they relate to one another.
How the model learns subjects, objects, and their links
The patent describes a three-step training pipeline for what Samsung calls a "foundation model" for scene graph generation (a technique that maps images as networks of labeled objects and the relationships connecting them).
- Step 1, Collect training data: The system gathers a large set of images paired with text captions that describe what's happening in each one.
- Step 2, Build a richer dataset: A large language model (an AI trained on text that already understands real-world relationships, like "a dog fetches a ball") is used to extract semantic knowledge (meaning-based understanding) and layer it onto the image-caption pairs. This turns a simple description into a structured set of entities (subjects and objects) and the relations between them.
- Step 3, Train the vision model: A separate machine learning model is trained on this enriched dataset so it can look at a new, unseen image and automatically identify entities and their relationships without needing a caption.
The claim is intentionally broad: it covers any electronic device doing this training process, and the model output applies to any input image. The patent does not limit itself to a specific type of scene, object category, or hardware.
… preparing, using the at least one processing device, a training dataset based on the training images, the captions, and semantic knowledge of a large language model.
Translation: Samsung uses AI language models to help the software understand the context of the images it is analyzing.
What this means for Samsung's camera and search features
Scene graph generation is one of those capabilities that sits underneath a lot of useful features: advanced photo search, automatic image captioning for accessibility, content moderation, and augmented reality overlays that understand context. If Samsung can build a strong foundation model for this task, it becomes a building block across its Galaxy devices and its broader AI software stack.
What makes this filing worth watching is the scope of claim 1. It covers the entire training method at a high level of abstraction, meaning any system that trains a model to find entities and their relations in images using language model knowledge could fall inside its boundary. Samsung's AI camera work is part of a busy field of latest Big Tech patents targeting computer vision as the next frontier for on-device AI.
This is the 98th Samsung filing we've tracked since May in our camera sensor push watchlist, building on work like a shorter signal path and a folded zoom design.
Claim 1 is written at a very high level of abstraction. It covers: get images and captions, use a large language model to prepare a dataset, train a model to find entities and relations. That's essentially a description of a broad class of scene-understanding research, not a specific technical invention. At that breadth, the patent examiner will almost certainly push back with prior art, because scene graph generation using language model supervision has been an active area of academic research for several years.
If Samsung narrows the claims during examination to something more specific, like a particular way of extracting relations from captions or a specific model architecture, the patent becomes more defensible but also less powerful as a blocking tool. As written, the first claim reads like a land-grab in territory that may already be well-mapped.
That said, the underlying engineering goal is real and commercially important. Getting a phone to understand a scene rather than just label objects in it is a meaningful step up. The question is whether Samsung's specific approach is novel enough to survive scrutiny.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
5 drawing sheets from US 2026/0253382 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →