Nvidia · Filed Mar 26, 2025 · Published Oct 1, 2026

Nvidia Patents a Fix for the Math That Blocked Deeper AI Compression

Compressing an AI model to run faster hits a wall whenever the model needs to add or subtract numbers, and Nvidia's new patent describes a way to knock that wall down.

A training framework processes a training data set to create a trained neural network, which then processes new data to produce a result. Drawing from patent filing US 2026/0300712 A1.
A training framework processes a training data set to create a trained neural network, which then processes new data to produce a result.
See all 8 drawings from this filing ↓
Publication number US 2026/0300712 A1
Applicant NVIDIA Corporation
Filing date Mar 26, 2025
Publication date Oct 1, 2026
Inventors Steve Dai, Muya Chang, Rangharajan Venkatesan
US classification 706/25
Status when we published Waiting for an examiner (Apr 16, 2025)
Document 33 claims

How Nvidia's shared-scale trick compresses more of an AI model

Imagine you're packing a suitcase and you can shrink most of your clothes, but a few bulky items refuse to compress no matter what. Running a modern AI model on a chip is similar: engineers can shrink most of its calculations to save time and power, but certain basic math steps, like adding two numbers together, have always been stubbornly resistant to that shrinking process.

The reason is simple: to add two numbers, they both need to be measured on the same scale. Current compression tools don't guarantee that, so engineers have had to leave those addition and subtraction steps uncompressed. Nvidia's patent describes a system that solves this by assigning a shared measuring scale to both numbers before the operation runs, so the math still works out after compression.

The payoff is that more of an AI model can be compressed end-to-end, which means the same AI could run faster, on less power, or on cheaper hardware, without retraining the model from scratch.

From the filing · CLAIM 1
determining one or more scale factors to be used for an element-wise operation of a neural network; scaling operands to the element-wise operation by the one or more scale factors to form scaled operands; …

Translation: The system calculates specific scaling numbers to adjust the data inputs before running them through the network math.

How the shared scale factors survive addition and subtraction

Neural networks are made up of many layers of calculations. Engineers speed them up through a process called quantization, essentially rounding numbers to a coarser scale (think switching from a ruler with millimeter markings to one with only centimeter markings). This shrinks the data each calculation touches, so the chip can work faster and use less energy.

The catch is that quantization assigns each number its own personal scaling factor (a multiplier that records how much it was shrunk). When you try to add or subtract two numbers that were shrunk by different factors, you can't just add the compressed versions directly, the scales are incompatible, like trying to add inches to centimeters without converting first. So these element-wise operations (additions and subtractions scattered throughout a neural network) have historically been left uncompressed, acting as bottlenecks.

Nvidia's approach introduces a shared scale factor applied to both operands (the inputs to an operation) before the addition or subtraction runs. Because both inputs are now on the same scale, the compressed math is valid. The output is also scaled, so downstream layers can continue working in the compressed format.

  • Determine one or more scale factors for an element-wise operation
  • Apply those factors to both inputs, producing scaled operands
  • Perform the addition or subtraction on the scaled operands to get a scaled output

This lets quantization propagate through parts of a network that were previously off-limits, pushing compression deeper into the model.

From the filing · THE ABSTRACT
… the degree to which a neural network can be quantized is limited because current scaling methods, which do not ensure consistent scaling for operands, cannot be applied to the commonly used element-wise operations (e.g. addition and subtraction).

Translation: Standard compression techniques fail on basic math operations like addition because they cannot keep the data scales matched up.

What deeper AI compression means for chips and real products

Every AI model that runs on a phone, a car's onboard computer, or a data-center chip is a balancing act between accuracy and efficiency. The more you compress a model, the faster and cheaper it runs, but today's compression tools stall out at those addition and subtraction steps, forcing hardware designers to leave headroom they'd rather not waste.

If this approach works as described, it means AI models could be compressed more aggressively without retraining, which is a significant cost in time and compute. For you, that could eventually translate to AI features that run locally on your device instead of in the cloud, respond faster, or drain your battery less. Nvidia's run of neural-network compression filings suggests this is a focused engineering priority, not a one-off idea.

Nvidia's 48th filing we've tracked since July in the AI chip wars continues a push to move work closer to the network, following patents on formatting data in the connector and offloading sending tasks to the GPU.

Editorial take

The cost of this approach is precision. When two numbers get forced onto the same measuring scale before being added together, neither one gets the scale that would represent it most accurately. That rounding error is the price of making compression possible at all.

For most practical uses, that price is probably acceptable. Decades of research show that AI models absorb small rounding errors remarkably well, and a bottleneck that prevents compression entirely is almost always a worse outcome than a modest accuracy dip that enables it.

What the patent leaves open is how much accuracy slips when this shared-scale approach gets applied across many layers of a model at once. That is an empirical question only real-world testing can answer, and it is where the engineering either holds up or falls apart.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

8 drawing sheets from US 2026/0300712 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.