Sony · Filed Mar 31, 2025 · Published Oct 1, 2026

Sony Patents a Single AI Vision System That Learns Multiple Visual Tasks at Once

Training a separate AI for every visual task, detecting objects, labeling scenes, measuring depth, is expensive and slow. Sony's new patent describes a way to build one shared visual brain that handles all of them together.

A single AI vision system uses a pretrained Vision Transformer and an adapter to perform various visual tasks like detection, segmentation, classification, and pose estimation. Drawing from patent filing US 2026/0301395 A1.
A single AI vision system uses a pretrained Vision Transformer and an adapter to perform various visual tasks like detection, segmentation, classification, and pose estimation.
See all 3 drawings from this filing ↓
Publication number US 2026/0301395 A1
Applicant Sony Group Corporation
Filing date Mar 31, 2025
Publication date Oct 1, 2026
Inventors Lingjuan LYU, Weiming ZHUANG, Zhizhong LI, Chen CHEN, Sina SAJADMANESH, Jingtao LI, Jiabo HUANG, Vivek SHARMA, Peter H. STONE, Michael S. SPRANGER
US classification 382/157
Examiner ZAK, JACQUELINE ROSE (Art Unit 2666)
Status when we published Waiting for an examiner (Apr 15, 2025)
Document 17 claims

What Sony's multi-task vision AI actually does

You're building a camera system that needs to recognize faces, read street signs, and separate objects from backgrounds, all at the same time. Right now, most AI setups require a separate model trained for each job, which costs time, money, and a lot of computing power.

Sony's patent describes a different approach: take a large, already-trained AI that's good at "seeing" (called a Vision Transformer), freeze it so it doesn't change, and attach a small trainable add-on called an adapter. That adapter learns to translate what the big AI sees into outputs for many different tasks simultaneously.

The result is one compact system that can be trained to do multiple vision jobs at once, rather than building and maintaining a fleet of separate AIs. It's designed to be more efficient to train and potentially easier to deploy in real products like cameras, robots, or autonomous vehicles.

From the filing · CLAIM 1
a frozen pre-trained vision transformer ‘ViT’; and a trainable adapter configured to interact with tokens of the ViT for successive interaction blocks of the ViT, by use of an injector and extractor; …

Translation: It combines a locked base vision model with a flexible add-on component that modifies the data flow.

How the adapter talks to the frozen vision model

The core of the system is a Vision Transformer (ViT), a large neural network pre-trained on enormous amounts of image data. Think of the ViT as a highly experienced visual analyst who already knows how to process images deeply. In Sony's design, this ViT is "frozen," meaning its learned knowledge is locked in place and not modified further.

Attached to the ViT is a trainable adapter, a lightweight module with two working parts: an injector and an extractor. The injector feeds task-relevant information into the ViT's processing pipeline at each layer (called a "block"), while the extractor pulls out useful signals at each stage. Together they let the frozen ViT's representations be steered toward whatever tasks are needed, without altering the ViT itself.

The adapter's output feeds into a set of decoders, each specialized for a different vision task, such as semantic segmentation (labeling every pixel in an image by what it belongs to), object detection, or depth estimation. All of these decoders are trained together in a process called multi-task learning, where the adapter and decoders learn simultaneously by sharing information across tasks.

  • Frozen ViT backbone: preserves expensive pre-training, no retraining cost
  • Injector/extractor adapter: steers the ViT's attention toward multiple tasks
  • Multiple decoders: each handles one visual task, trained together in parallel
From the filing · THE ABSTRACT
… the trainable adapter is trained in conjunction with training a plurality of different decoders comprising a first set of decoders, respectively coupled to the backbone for respective tasks, to produce the completed vision foundation model based on multi-task learning.

Translation: The system learns many different visual jobs at once by pairing the core model with multiple specialized task decoders.

What this means for cameras and computer vision devices

Training large AI vision models from scratch is expensive enough that only a handful of companies can do it regularly. Sony's approach cuts that cost by reusing a frozen pre-trained model and only training the small adapter and decoders on top. That makes capable multi-task vision AI more accessible and faster to update when new tasks are needed.

For products like smart cameras, robots, and autonomous vehicles, the ability to run one shared visual model across many tasks, rather than several separate models, reduces hardware demands and simplifies deployment. a growing pile of Sony AI vision filings suggests this is part of a broader push to bring foundation-model thinking into Sony's imaging and sensor hardware lines.

Sony's 11th filing in the AI vision patents we've tracked since May follows work like one on ghost image removal and one on free-angle scene freezing.

Editorial take

Building a separate visual recognition system for every task a product needs, reading faces, spotting objects, parsing text, costs engineering teams months of work and serious compute budget each time. Across a full product line, that repeated rebuilding becomes one of the largest drains on AI development.

Sony's patent attacks that cost directly by training one small shared layer that sits atop a single large visual core, letting it serve many tasks without rebuilding from scratch each time. The approach fits the actual scale of the problem.

What the patent describes is a structure rather than a proven result, so the real efficiency gains remain to be measured. But the underlying bet looks sound: the waste is real, and consolidating around a shared adaptable core is a proportionate response to it.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

3 drawing sheets from US 2026/0301395 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.