Qualcomm · Filed Mar 25, 2025 · Published Oct 1, 2026 · verified — real USPTO data

Qualcomm Patents AI That Reads and Describes Real-World Scenes in Three Dimensions

Teaching a computer to understand what's around it in three dimensions requires enormous amounts of carefully labeled data, and Qualcomm thinks it has a smarter way to cut that cost without cutting the results.

A foundation model processes a LiDAR point cloud to achieve 3D scene understanding through multi-modal encoding and transformer decoding. Drawing from patent filing US 2026/0301377 A1.
A foundation model processes a LiDAR point cloud to achieve 3D scene understanding through multi-modal encoding and transformer decoding.
See all 8 drawings from this filing ↓
Publication number US 2026/0301377 A1
Applicant QUALCOMM Incorporated
Filing date Mar 25, 2025
Publication date Oct 1, 2026
Inventors Venkatraman Narayanan, Varun Ravi Kumar, Senthil Kumar Yogamani
CPC classification 382/103
Grant likelihood Medium
Examiner TC 4100, DOCKET (Art Unit 4100)
Status Docketed New Case - Ready for Examination (Apr 9, 2025)
Document 20 claims

What Qualcomm's 3D scene-understanding AI actually does

Ever tried to describe a crowded parking lot to someone over the phone? Now think about how hard it is for a machine to do that, in real time, from raw sensor readings. That's essentially the problem Qualcomm is trying to solve here.

The patent describes an AI model that learns to understand 3D environments in two steps. First, it trains on a smaller batch of data that humans have carefully labeled, think a self-driving car dataset where every pedestrian, cone, and car has been tagged by hand. Then, in a second pass, it keeps learning from much larger pools of unlabeled data, sensor recordings where nobody bothered to tag anything, which are far cheaper and easier to collect.

The system can also take hints from humans on the fly when it encounters an object it doesn't recognize. You get an AI that starts with a solid foundation and keeps getting better without needing a small army of data labelers.

From the filing · CLAIM 1
… training a foundation model on three-dimensional scene understanding based on the labeled three-dimensional sensor data a first training stage, wherein the foundation model comprises a multimodal large language model (LLM) encoder, a fusion transformer, and a transformer decoder; …

Translation: Training the AI system using labeled data with a combination of specialized language and translation components.

How the two-stage training and fusion model work together

The core of the patent is a foundation model (a large, general-purpose AI that can be applied to many tasks) built specifically for interpreting 3D sensor data, the kind that comes from LiDAR scanners or depth cameras used in autonomous vehicles and robots.

The model has three main components working in sequence:

  • Multimodal LLM encoder: A large language model reads the 3D sensor data and generates plain-text descriptions of the objects it detects, essentially writing captions like "a person standing near a door" from point-cloud data (a 3D map made of millions of tiny measured points).
  • Fusion transformer: A neural network layer combines those text descriptions with the raw 3D data into a single unified representation, letting language and geometry inform each other.
  • Transformer decoder: The final stage uses that combined representation to actually perform perception tasks, detecting objects, segmenting scenes, or classifying what's around the sensor.

Training happens in two stages. Stage one uses supervised learning on labeled data (examples where a human has already identified what each object is). Stage two switches to self-supervised learning (the model generates its own training signal from unlabeled recordings). The second stage is what makes the approach scalable, unlabeled sensor data is abundant and cheap compared to hand-annotated datasets.

An online learning component also lets the model update in real time when a human flags an object the model couldn't classify, which is a practical bridge between lab training and real-world deployment.

From the filing · THE ABSTRACT
A two-stage training approach is utilized that combines supervised learning on labeled three-dimensional sensor data in the first stage and self-supervised learning on large, unlabeled sensor datasets in the second stage.

Translation: The system learns first from tagged data and then improves itself using massive amounts of raw data.

What this means for autonomous vehicles and robotics

For autonomous vehicles and robotics, 3D scene understanding is one of the hardest and most expensive problems to solve. Collecting and labeling enough training data is a bottleneck that slows down nearly every team working in this space. A system that can extend its own learning from raw, unlabeled sensor logs, without a constant stream of human annotation, could significantly lower that barrier.

For you as a consumer, the downstream effect would be systems that handle unexpected real-world situations more reliably: a delivery robot that recognizes an unusual obstacle, or a car that correctly identifies an object it's never seen in training. The patent also sits squarely in Qualcomm's hardware business, the company makes the chips that power many edge AI systems, so a capable 3D understanding model optimized for their architecture has direct commercial value for their customers.

Qualcomm's 20th filing we've tracked in Enterprise AI since May adds to a run that includes one on running leaner AI on phones and one on cutting video data AI reads.

Editorial take

The two-stage training design is a reasonable answer to a real problem, but the tradeoffs deserve some scrutiny. Self-supervised learning on unlabeled 3D data is notoriously harder to stabilize than its 2D image equivalent, the model has to generate useful learning signals from sparse, noisy point clouds, and there's no guarantee the second stage actually improves on the first rather than drifting away from it.

The online learning component, where the model updates from human feedback during deployment, is the riskiest piece. Real-time model updates in safety-critical systems like autonomous vehicles raise genuine questions about consistency and verification that the patent doesn't address. That's understandable for a patent filing, but it's the gap between the idea and a shippable product.

The fusion of language model outputs with 3D sensor data is the most interesting design choice here. Borrowing the representational power of a large language model to describe geometry is an unusual move, and it could pay off if the text descriptions actually capture semantic structure that raw point clouds miss. Whether that benefit outweighs the added computational cost of running an LLM in the perception pipeline is a question this filing leaves open.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

8 drawing sheets from US 2026/0301377 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.