Sony Patents a Single AI Vision System That Learns Multiple Visual Tasks at Once
Training a separate AI for every visual task, detecting objects, labeling scenes, measuring depth, is expensive and slow. Sony's new patent describes a way to build one shared visual brain that handles all of them together.
What Sony's multi-task vision AI actually does
You're building a camera system that needs to recognize faces, read street signs, and separate objects from backgrounds, all at the same time. Right now, most AI setups require a separate model trained for each job, which costs time, money, and a lot of computing power.
Sony's patent describes a different approach: take a large, already-trained AI that's good at "seeing" (called a Vision Transformer), freeze it so it doesn't change, and attach a small trainable add-on called an adapter. That adapter learns to translate what the big AI sees into outputs for many different tasks simultaneously.
The result is one compact system that can be trained to do multiple vision jobs at once, rather than building and maintaining a fleet of separate AIs. It's designed to be more efficient to train and potentially easier to deploy in real products like cameras, robots, or autonomous vehicles.
a frozen pre-trained vision transformer ‘ViT’; and a trainable adapter configured to interact with tokens of the ViT for successive interaction blocks of the ViT, by use of an injector and extractor; …
Translation: It combines a locked base vision model with a flexible add-on component that modifies the data flow.
How the adapter talks to the frozen vision model
The core of the system is a Vision Transformer (ViT), a large neural network pre-trained on enormous amounts of image data. Think of the ViT as a highly experienced visual analyst who already knows how to process images deeply. In Sony's design, this ViT is "frozen," meaning its learned knowledge is locked in place and not modified further.
Attached to the ViT is a trainable adapter, a lightweight module with two working parts: an injector and an extractor. The injector feeds task-relevant information into the ViT's processing pipeline at each layer (called a "block"), while the extractor pulls out useful signals at each stage. Together they let the frozen ViT's representations be steered toward whatever tasks are needed, without altering the ViT itself.
The adapter's output feeds into a set of decoders, each specialized for a different vision task, such as semantic segmentation (labeling every pixel in an image by what it belongs to), object detection, or depth estimation. All of these decoders are trained together in a process called multi-task learning, where the adapter and decoders learn simultaneously by sharing information across tasks.
- Frozen ViT backbone: preserves expensive pre-training, no retraining cost
- Injector/extractor adapter: steers the ViT's attention toward multiple tasks
- Multiple decoders: each handles one visual task, trained together in parallel
… the trainable adapter is trained in conjunction with training a plurality of different decoders comprising a first set of decoders, respectively coupled to the backbone for respective tasks, to produce the completed vision foundation model based on multi-task learning.
Translation: The system learns many different visual jobs at once by pairing the core model with multiple specialized task decoders.
What this means for cameras and computer vision devices
Training large AI vision models from scratch is expensive enough that only a handful of companies can do it regularly. Sony's approach cuts that cost by reusing a frozen pre-trained model and only training the small adapter and decoders on top. That makes capable multi-task vision AI more accessible and faster to update when new tasks are needed.
For products like smart cameras, robots, and autonomous vehicles, the ability to run one shared visual model across many tasks, rather than several separate models, reduces hardware demands and simplifies deployment. a growing pile of Sony AI vision filings suggests this is part of a broader push to bring foundation-model thinking into Sony's imaging and sensor hardware lines.
Sony's 11th filing in the AI vision patents we've tracked since May follows work like one on ghost image removal and one on free-angle scene freezing.
Building a separate visual recognition system for every task a product needs, reading faces, spotting objects, parsing text, costs engineering teams months of work and serious compute budget each time. Across a full product line, that repeated rebuilding becomes one of the largest drains on AI development.
Sony's patent attacks that cost directly by training one small shared layer that sits atop a single large visual core, letting it serve many tasks without rebuilding from scratch each time. The approach fits the actual scale of the problem.
What the patent describes is a structure rather than a proven result, so the real efficiency gains remain to be measured. But the underlying bet looks sound: the waste is real, and consolidating around a shared adaptable core is a proportionate response to it.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
3 drawing sheets from US 2026/0301395 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in