Amazon Patents a Way to Compress AI Math by Treating Tricky Numbers Differently
Running a large AI model is expensive partly because the numbers inside it take up a lot of memory. Amazon has filed a patent for a system that shrinks those numbers more aggressively by first sorting them into two buckets: the well-behaved ones and the messy outliers.
What Amazon's two-tier AI compression actually does
You're watching a video stream and it looks sharp the whole way through, even on a cheap device. That quality comes partly from AI models running in the background, and those models are memory-hungry. Cutting their memory use is one of the biggest engineering problems in the industry right now.
Amazon's patent describes a system that compresses the numbers inside a neural network more efficiently by treating different numbers differently. Instead of applying one-size-fits-all compression across an entire layer of the AI, the system flags rows of numbers that have unusual, hard-to-compress values, and keeps those at higher precision. The ordinary rows get squeezed down to a smaller format that takes up less memory and runs faster.
The result is a single compressed layer that mixes two levels of detail, handling the awkward data carefully while shrinking everything else as aggressively as possible. The goal is accuracy close to the uncompressed original, at a fraction of the memory cost.
… determining, based on the plurality of activation values, a set of outlier rows; determining, based on the plurality of activation values, a set of non-outlier rows; …
Translation: The system splits the data into typical values and problematic outliers to handle them separately.
How the system splits rows before shrinking the numbers
Neural networks process data as large grids of numbers called activation matrices. These matrices move between layers of the network during inference (when the AI is actually running and answering questions). Storing and moving those matrices is expensive in both memory and energy.
Quantization is the standard fix: you replace high-precision floating-point numbers (which might use 16 or 32 bits each) with lower-precision integers (8, 4, or even 2 bits). Smaller numbers mean smaller memory footprint and faster arithmetic. The problem is that neural networks often have outlier values, numbers that are much larger or more extreme than the rest. Aggressive compression distorts these outliers badly, and that distortion cascades through the rest of the network as errors.
Amazon's approach works like this:
- The system computes the activation matrix for a given layer and inspects its rows of values.
- Rows that contain outlier values are flagged as a separate set.
- Two different quantization algorithms are chosen: a higher-bitwidth one for the outlier rows (preserving their detail) and a lower-bitwidth one for the ordinary rows (compressing them hard).
- Both are applied within the same layer, producing a single mixed-precision matrix that the next layer then uses for its own computation.
The hardware target is a neural network accelerator that uses both fast on-chip SRAM and larger off-chip DRAM, which is the memory architecture common in AI inference chips like those Amazon builds for AWS.
A first quantization algorithm may be selected for the set of outlier rows. A second quantization algorithm may be selected for the set of non-outlier rows.
Translation: Different compression methods are applied depending on how tricky the numbers are to process.
What this means for the cost of running AI at scale
For anyone paying to run AI services, memory bandwidth is one of the biggest line items. Compressing activation matrices harder means you can fit larger models onto the same chip, or run the same model faster, or both. That translates into lower cost per query for cloud providers and potentially faster, cheaper on-device AI for consumer products.
The specific challenge this addresses, outlier values wrecking uniform compression, is a well-known pain point in deploying large language models. the pattern in Amazon's AI hardware filings suggests the company is building toward tighter control of the full inference stack, from the chip architecture up through the software that schedules how numbers are stored and moved. A system that handles outliers inside individual layers, rather than at the whole-model level, gives engineers finer control over the accuracy-versus-speed tradeoff.
Amazon's tenth filing we've tracked in the AI chip wars since May follows earlier applications like one on surviving server crashes and one splitting quantum and regular work.
The design asks the chip to stop and sort every layer of data into two piles before compressing it, and that sorting step burns time and energy on every single pass through the network. Someone also has to set the rule for what counts as an outlier, and a badly chosen threshold means the whole system degrades without obvious warning.
For large language models, where a small number of extreme values can wreck the accuracy of the entire compressed output, that cost probably earns its keep. The quality gain at the same storage size is real enough to justify the extra work.
The harder question is permanence. Building this logic into the chip itself means the sorting behavior cannot be easily adjusted once the hardware ships, so if the threshold proves wrong for a new model or use case, there is no software patch that fully fixes it.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
7 drawing sheets from US 2026/0300701 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in