Sony Patent Blends Real and Synthetic Depth Data to Sharpen AI Vision
Training an AI to see in 3D requires enormous amounts of labeled data, and gathering it in the real world is slow and expensive. Sony's new patent cuts that cost by manufacturing the data itself.
What Sony's synthetic depth training data actually does
Imagine teaching a self-driving car or a robot arm to judge distances. The AI needs thousands of examples showing exactly how far away each object is. Collecting all of that in the real world takes serious time and money, and some scenarios are hard or impossible to capture safely.
Sony's approach is to take a real depth image of a scene (a 3D map of distances, captured by a sensor) and then drop in a computer-generated object on top of it. The result looks like the fake object was always part of the real scene. That blended image becomes training data.
This lets Sony (or anyone using this method) generate huge libraries of training examples cheaply, covering rare objects or tricky situations that almost never appear in ordinary footage. The AI gets practice spotting things it may rarely encounter in the wild, without anyone having to physically stage those situations.
How Sony merges real scenes with computer-generated objects
The patent describes a two-step pipeline. First, a synthetic depth image of a computer-generated object is created. A depth image is not a regular photo; it encodes distance rather than color, so each pixel carries a number representing how far away that point in space is from the sensor.
Second, that synthetic depth image is merged with a real depth image captured from an actual physical scene. The merger produces a composite image in which the fake object appears to sit naturally inside the real environment, complete with plausible distance values.
The patent also mentions extending the same technique to confidence images (which track how certain the sensor is about each distance reading) and segmentation images (which label which pixels belong to which object). That means the training signal covers not just where things are but how reliably the sensor can see them.
The resulting blended images are fed directly into an artificial neural network as labeled training data, bypassing the need to physically place real objects in real environments and manually annotate them.
What this means for depth-sensing cameras and AI vision
Depth-sensing hardware (the kind found in robotics, AR headsets, and advanced cameras) is only as useful as the AI model interpreting its output. Getting enough labeled training data has always been a bottleneck, especially for uncommon objects or unusual lighting conditions. This method lets engineers generate that data on demand, which can translate to more accurate depth perception in your devices without proportionally higher development costs.
Sony Semiconductor Solutions makes image sensors used across many industries, so a more efficient training pipeline here has practical reach. If this technique tightens the accuracy of depth models, the downstream effect could show up in things like better autofocus, more reliable gesture control, or safer industrial robots.
This is a sensible, workmanlike approach to a real problem in computer-vision development. Synthetic data augmentation is not a new idea, but applying it specifically to depth and confidence images (rather than just RGB photos) is the concrete contribution here. It is not flashy, but for any team building depth-sensing AI it is a meaningful time-saver.
The drawings
15 drawing sheets from US 2026/0220874 A1 · click any drawing to enlarge
Which company should we read for you?
We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.
Get one Big Tech patent every Sunday
Plain English, intelligent commentary, no hype. Free.
Editorial commentary on a publicly published patent application. Not legal advice.