Nvidia Patents a Fix for the Math That Blocked Deeper AI Compression
Compressing an AI model to run faster hits a wall whenever the model needs to add or subtract numbers, and Nvidia's new patent describes a way to knock that wall down.
How Nvidia's shared-scale trick compresses more of an AI model
Imagine you're packing a suitcase and you can shrink most of your clothes, but a few bulky items refuse to compress no matter what. Running a modern AI model on a chip is similar: engineers can shrink most of its calculations to save time and power, but certain basic math steps, like adding two numbers together, have always been stubbornly resistant to that shrinking process.
The reason is simple: to add two numbers, they both need to be measured on the same scale. Current compression tools don't guarantee that, so engineers have had to leave those addition and subtraction steps uncompressed. Nvidia's patent describes a system that solves this by assigning a shared measuring scale to both numbers before the operation runs, so the math still works out after compression.
The payoff is that more of an AI model can be compressed end-to-end, which means the same AI could run faster, on less power, or on cheaper hardware, without retraining the model from scratch.
determining one or more scale factors to be used for an element-wise operation of a neural network; scaling operands to the element-wise operation by the one or more scale factors to form scaled operands; …
Translation: The system calculates specific scaling numbers to adjust the data inputs before running them through the network math.
How the shared scale factors survive addition and subtraction
Neural networks are made up of many layers of calculations. Engineers speed them up through a process called quantization, essentially rounding numbers to a coarser scale (think switching from a ruler with millimeter markings to one with only centimeter markings). This shrinks the data each calculation touches, so the chip can work faster and use less energy.
The catch is that quantization assigns each number its own personal scaling factor (a multiplier that records how much it was shrunk). When you try to add or subtract two numbers that were shrunk by different factors, you can't just add the compressed versions directly, the scales are incompatible, like trying to add inches to centimeters without converting first. So these element-wise operations (additions and subtractions scattered throughout a neural network) have historically been left uncompressed, acting as bottlenecks.
Nvidia's approach introduces a shared scale factor applied to both operands (the inputs to an operation) before the addition or subtraction runs. Because both inputs are now on the same scale, the compressed math is valid. The output is also scaled, so downstream layers can continue working in the compressed format.
- Determine one or more scale factors for an element-wise operation
- Apply those factors to both inputs, producing scaled operands
- Perform the addition or subtraction on the scaled operands to get a scaled output
This lets quantization propagate through parts of a network that were previously off-limits, pushing compression deeper into the model.
… the degree to which a neural network can be quantized is limited because current scaling methods, which do not ensure consistent scaling for operands, cannot be applied to the commonly used element-wise operations (e.g. addition and subtraction).
Translation: Standard compression techniques fail on basic math operations like addition because they cannot keep the data scales matched up.
What deeper AI compression means for chips and real products
Every AI model that runs on a phone, a car's onboard computer, or a data-center chip is a balancing act between accuracy and efficiency. The more you compress a model, the faster and cheaper it runs, but today's compression tools stall out at those addition and subtraction steps, forcing hardware designers to leave headroom they'd rather not waste.
If this approach works as described, it means AI models could be compressed more aggressively without retraining, which is a significant cost in time and compute. For you, that could eventually translate to AI features that run locally on your device instead of in the cloud, respond faster, or drain your battery less. Nvidia's run of neural-network compression filings suggests this is a focused engineering priority, not a one-off idea.
Nvidia's 48th filing we've tracked since July in the AI chip wars continues a push to move work closer to the network, following patents on formatting data in the connector and offloading sending tasks to the GPU.
The cost of this approach is precision. When two numbers get forced onto the same measuring scale before being added together, neither one gets the scale that would represent it most accurately. That rounding error is the price of making compression possible at all.
For most practical uses, that price is probably acceptable. Decades of research show that AI models absorb small rounding errors remarkably well, and a bottleneck that prevents compression entirely is almost always a worse outcome than a modest accuracy dip that enables it.
What the patent leaves open is how much accuracy slips when this shared-scale approach gets applied across many layers of a model at once. That is an empirical question only real-world testing can answer, and it is where the engineering either holds up or falls apart.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
8 drawing sheets from US 2026/0300712 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in