Sony Patents a Dual-AI Method for Cutting Subjects Out of Photos
Getting a clean cutout of a person or object in a photo is harder than it looks, and Sony thinks two AI models working together can do it better than one alone.
What Sony's two-model silhouette system actually does
You're trying to scan a friend in full 3D, and the software keeps including bits of the wall behind them or clipping off their shoulder. The problem is that no single method of figuring out where a person ends and the background begins is perfect in every situation.
Sony's patent describes a system that runs two separate approaches at once and then blends their results. One AI model looks at both the photo and a plain background image to spot what changed. A second method identifies objects by recognizing what they are. By merging these two outputs, the system can catch mistakes that either approach would make on its own.
The target use case is building 3D models from photos, where a sloppy cutout leaves messy edges on the final figure. Cleaner silhouettes mean cleaner 3D scans, whether that's for games, film production, or product photography.
… merge processing unit configured to merge a first silhouette image output from a first learning model that takes, as input, a captured image in which a subject and a background appear and a background image in which the background appears …
Translation: This component combines cutout results from an AI that compares a photo against a clean background shot.
How the two silhouette outputs get merged into one
The patent centers on a merge processing unit that combines two independently generated silhouette images into one final result.
The first silhouette comes from a learned AI model (trained on image data) that takes two inputs: the actual photo of a subject against a background, and a reference photo of that same background without the subject. Because it can compare the two, it gets a strong signal about which pixels belong to the person or object and which belong to the scenery behind them.
The second silhouette is generated through object recognition, meaning the system looks at the photo and tries to identify and outline known categories of things (people, chairs, products, etc.). This approach doesn't need a clean background reference shot, but it can struggle with unusual poses or lighting.
- Model 1 strength: precise edge detection when a background reference exists
- Model 2 strength: object-aware, works even without a background-only shot
- Merge unit: combines both outputs to cover each model's blind spots
The merged result is meant to feed into 3D model generation pipelines, where each camera angle needs a reliable silhouette to reconstruct the subject's shape in three dimensions.
… an image processing device that extracts a silhouette of a subject from a captured image for use in generating a 3D model …
Translation: The system isolates people or objects in pictures so they can be turned into three dimensional objects.
What this means for 3D scanning and photo editing
For anyone building 3D scans, whether for movies, games, e-commerce product pages, or virtual try-on, cleaner subject cutouts directly mean less manual cleanup work. Right now, artists often spend time fixing jagged or incomplete edges by hand. A system that gets the outline right automatically saves real production hours.
On the consumer side, this kind of technology feeds into camera apps, video call background removal, and augmented reality features. Sony's track record in image-processing patents suggests the company sees this as infrastructure that could run across its camera hardware, PlayStation platform, and professional imaging tools. If the merge approach genuinely outperforms single-method cutouts, it becomes a useful building block across a wide range of products.
Sony's 17th filing we've tracked since July in the AI photo editing race, following one fixing extreme lighting and one keeping video frames consistent, shows how broadly the company is applying AI to image work.
Claim 1 is written broadly. It covers any device that merges a silhouette from a background-comparison AI model with a silhouette from an object-recognition method. That scope doesn't specify the architecture of either model, the format of the merge, or the end application. A claim that wide could, if granted, reach a large share of multi-model subject-extraction systems.
In practice, that breadth is also the claim's weakness. Background-subtraction plus object detection is a well-established combination in computer vision research, and the patent office will look hard at prior art. The novel hook here seems to be the specific pairing of a model that takes a background reference image as input with a separate object-recognition pass, but that distinction may not be enough to survive a thorough prior-art search.
For readers, the honest read is that this is a solid engineering approach to a real problem, but the patent is more interesting as a signal of Sony's direction in 3D capture than as a locked-down exclusive technique.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
15 drawing sheets from US 2026/0301123 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in