Waymo · Filed Mar 30, 2026 · Published Oct 1, 2026

Waymo Patents a Prediction System That Shows Self-Driving Cars What Comes Next

Waymo is teaching a single neural network to absorb every sensor a self-driving car has, all at once, and then predict what the road will look like moments from now. That internal crystal ball is what lets the car plan moves before events fully unfold.

A self-driving car's on-board system uses sensor data and trained model parameters for prediction, planning, and user interaction. Drawing from patent filing US 2026/0296454 A1.
A self-driving car's on-board system uses sensor data and trained model parameters for prediction, planning, and user interaction.
See all 11 drawings from this filing ↓
Publication number US 2026/0296454 A1
Applicant Waymo LLC
Filing date Mar 30, 2026
Publication date Oct 1, 2026
Inventors Chiyu Jiang, Harish Chandran
US classification 701/48
Status when we published Waiting for an examiner (May 21, 2026)
Parent application Claims priority from a provisional application 63780046 (filed 2025-03-28)
Document 20 claims

What Waymo's sensor-fusion world model actually does

Ever wondered how a self-driving car decides what to do half a second before something happens? It has to predict the future, not just react to the present.

Waymo's patent describes a system that takes data from all of a car's sensors at once, cameras, radar, lidar, and more, and feeds it into a single neural network as a stream of timed snapshots. The network has been trained on huge numbers of real driving situations, so it learns to guess what comes next: where other cars are heading, whether a pedestrian is about to step off the curb, how the scene will look a few moments from now.

The output is a prediction the car can act on. Instead of treating each sensor as its own separate problem, the system sees everything together, which gives it a more complete picture of the world around it.

From the filing · CLAIM 1
obtaining an input sequence of tokens characterizing a driving environment for a vehicle, wherein the input sequence of tokens comprises, for each of one or more time steps, a respective set of tokens for each of a plurality of modalities of data from a set of multiple modalities of data …

Translation: The system collects different types of environmental data over time to understand what is happening around the car.

How the network turns sensor feeds into future predictions

The patent describes what engineers call a world model: a neural network trained not just to classify what it sees, but to forecast how a scene will evolve over time.

Input arrives as a sequence of tokens (think of tokens as compact numerical summaries of a moment in time) drawn from multiple modalities (different sensor types: camera images, lidar point clouds, radar returns). For each time step in recent history, the network receives one set of tokens per sensor type, all stitched together into a single ordered list.

The network then processes that entire list to generate an output token sequence that represents a predicted future state. That prediction can serve several downstream tasks:

  • Estimating where nearby vehicles and pedestrians will be in the next few seconds
  • Validating the car's planned path against likely futures
  • Flagging scenarios where the environment is about to change rapidly

Critically, the network is trained end-to-end on example driving environments, meaning it learns the relationships between sensor types implicitly, without engineers hand-coding rules about how camera data should interact with lidar data. The architecture is described as multi-modal token processing, a design borrowed from large language models but applied to physical sensor streams instead of text.

From the filing · THE ABSTRACT
the multi-modal token processing neural network has been trained to process input token sequences characterizing current states of example driving environments to generate output token sequences characterizing predicted future states of the example driving environments …

Translation: The artificial intelligence learns from past driving situations to forecast what will happen next on the road.

What this means for self-driving car reliability

For passengers and pedestrians, a car that anticipates the near future is meaningfully safer than one that only reacts to the present. Reaction-only systems have to wait for an event to fully register across sensors before they can respond. A world model that forecasts what is about to happen gives the car a small but real head start.

For Waymo's engineering team, fusing all sensor types through a single trained network also reduces the number of hand-tuned rules that have to be maintained as hardware changes. If a new sensor type is added, it becomes another modality in the token stream rather than a separate processing pipeline. That kind of architectural flexibility matters when you are running a commercial robotaxi fleet and updating software continuously.

Google's 63rd filing we've tracked since May in the self-driving sensing race adds to a run that includes one grading its own decisions and one predicting road users' paths.

Editorial take

Routing every sensor through one shared network means the system succeeds or fails as a whole. If something goes wrong, there is no clean way to point at the camera section or the radar section and fix it in isolation, which makes failures harder to diagnose and harder to explain to regulators or passengers.

That opacity is the real cost of this design, and it is a serious one. A network this large also has to run fast enough that the car can act on what it sees, and the patent says nothing about how that speed requirement gets met.

The trade still reads as defensible for a company building toward full autonomy rather than a narrow product feature. A single model that has absorbed all the sensor data together will probably catch things that separate, siloed models would miss, and Waymo's history suggests they are willing to absorb short-term engineering pain for long-term architectural leverage.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

11 drawing sheets from US 2026/0296454 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.