Tesla · Filed Sep 8, 2025 · Published Oct 1, 2026

Tesla Patents a Per-Token Precision Trick to Speed Up Its Self-Driving AI

Tesla's newest patent tackles a problem every AI chip faces: the math that makes AI accurate is too slow for real-time driving, but the math that's fast enough tends to make mistakes. This filing describes a way to have both.

A self-driving car and a human with computing devices connect to a network and servers for AI processing. Drawing from patent filing US 2026/0299881 A1.
A self-driving car and a human with computing devices connect to a network and servers for AI processing.
See all 12 drawings from this filing ↓
Publication number US 2026/0299881 A1
Applicant Tesla, Inc.
Filing date Sep 8, 2025
Publication date Oct 1, 2026
Inventors Ritvik RAWAT, Valeriy ROTAN, Alex SINGH, Srihari SAMPATHKUMAR
US classification 706/21
Status when we published Waiting for an examiner (Sep 26, 2025)
Parent application Claims priority from a provisional application 63780053 (filed 2025-03-28)
Document 20 claims

How Tesla's self-driving chip balances speed and accuracy

Ever tried to cram a high-resolution photo into a tiny file without making it look terrible? That tradeoff, quality versus size, is basically what AI chips face every time they process data.

Your car's self-driving system processes sensor readings dozens of times per second. The most accurate calculations use big, precise numbers, but those take more time and power. Cheaper, smaller numbers are faster but less accurate. Tesla's patent describes a method where the chip automatically figures out, for each individual chunk of data, exactly how much precision that chunk needs, then converts it on the fly before doing the heavy math.

The result is that the AI can use fast, low-power hardware for most of the work, while keeping accuracy where it counts. That means the same chip can do more thinking in less time, which matters a lot when your car is deciding whether something in the road is a shadow or a rock.

From the filing · CLAIM 1
generating, by the processor, a token scale factor based on a range of values associated with the token; and executing, by the processor, a quantization operation on the token of the input data using the token scale factor …

Translation: It calculates a custom scale for each token to safely shrink its data size.

How the quantization layer rescales each token before compute

The patent describes a transformer architecture (the same class of AI model that powers large language models, adapted here for sensor data) that runs in mixed precision: different parts of the calculation use different numeric formats depending on what is needed.

The key mechanism is a dynamic quantization layer. "Quantization" means converting a high-precision floating-point number (think: a decimal carried out to many places) into a lower-precision integer (a rounded whole number). Tesla's system does this not once globally, but per token. A "token" here is a discrete chunk of input data, such as a patch of a camera image or a slice of a LiDAR scan.

For each token, the system:

  • Inspects the range of values inside that token
  • Calculates a scale factor tuned to that specific range
  • Converts the token from high-precision (FP32, or 32-bit floating point) down to low-precision (INT8, or 8-bit integer) using that scale factor
  • Passes the scaled-down token to fast integer multiply-accumulate (MAC) hardware, which is far more power-efficient than floating-point units

Because each token gets its own scale factor rather than sharing one global setting, the precision loss is minimized. The hardware described includes both INT8 MAC units (fast, efficient integer math) and FP32 SIMD processors (slower but fully precise, used where needed), letting the chip allocate the right tool to each part of the job.

From the filing · THE ABSTRACT
The machine-learning architecture implements a transformer that includes dynamic quantization layers that dynamically quantize tensors, per-token, for downstream operations …

Translation: The system adjusts the data precision token by token as it flows through the AI model.

What this means for Tesla's in-car AI chip performance

For Tesla, the constraint is real: the AI computer inside a car must process multiple camera feeds, radar, and other sensors in real time, all on a power budget that won't drain the battery or overheat the cabin. Getting more AI throughput without bigger, hotter chips is a genuine engineering priority, and per-token dynamic quantization is one of the more principled ways to do it.

For you as a driver, this kind of optimization is what allows the same physical hardware to handle increasingly capable AI models over time, potentially through software updates rather than new chips. the pattern in Tesla's AI silicon filings points toward the company wanting its vehicles to do more computation onboard rather than relying on cloud processing, and this patent fits that direction precisely.

Tesla's second filing we've tracked in the AI chip competition since October follows its method to tune in-car chips.

Editorial take

Claim 1 describes a processor that receives data, decides how precisely to represent it, generates a scaling factor for each individual chunk of that data, compresses it, and then runs a calculation. Nothing in that language mentions vehicles, cameras, or self-driving software. The claim covers the method itself, on any device, running any neural network.

That scope has real consequences. Any company building AI chips or software that dynamically adjusts numerical precision on a per-chunk basis before a matrix calculation would fall inside this claim if it grants as written.

Whether the claim survives examination depends entirely on whether this exact sequence already appears somewhere in prior research. If it does not, this patent would give Tesla meaningful leverage over a core technique in AI hardware well beyond the self-driving context where the work originated.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

12 drawing sheets from US 2026/0299881 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.