AMD Files Patent for Running Image AI on About a Third of the Memory
AMD's new patent application describes a way to run image-recognition AI by working through its data in small batches instead of all at once. The filing estimates that could cut the storage needed to about a third.
What AMD's split-up image AI patent does
Imagine you have to add up a giant pile of receipts, but your desk only fits a few at a time. You wouldn't spread them all out. You'd work through one stack, jot down a running total, push the finished stack aside, and grab the next.
That's the idea behind this AMD patent application. Image-recognition AI normally takes in every layer of a picture, such as its color information, all at once, and that eats a lot of memory. AMD's filing splits those layers into groups, works on one group at a time, and adds up the results as it goes. It also keeps the in-between results in a small, fast memory next to the chip instead of sending them out to the main memory.
The filing estimates this could cut the storage needed to about a third, and trips to main memory to about 8% of what they'd otherwise be. That matters most on phones and other devices short on power and space.
… subdividing the plurality of channels into groups of intermediate channels; executing a plurality of convolutions based on the intermediate channels to produce an intermediate output for each group of intermediate channels; and adding the intermediate output for each group of intermediate channels to generate output feature data.
Translation: The chip breaks image data into smaller pieces and processes them one at a time to save memory.
How the channel groups get added back up
A convolutional neural network (CNN, an AI design that spots patterns in images) normally takes in every channel of its input at once. A channel is one layer of the data, such as the color layers of a photo. Holding all of them while the math runs takes a lot of memory.
The main claim does this instead:
- Split the channels into groups (the filing calls them "intermediate channels").
- Run a set of convolutions (passes where a small grid of learned numbers slides across the data to find edges and textures) on one group, producing an intermediate output.
- Add each group's output into a running total to make the final output.
The description fills in the steps per group: a 1x1 filter, then a depthwise convolution (a separate filter for each channel) with a 3x3 or 5x5 grid, then another 1x1 filter. After each group, input data no longer needed is discarded. Groups can be picked randomly, in order, or by priority.
The second idea is a buffer, a small, fast memory in or near the processor, that holds intermediate results instead of main system memory. A third independent claim splits the whole network into subnetworks, one per group, and runs them in sequence. Example inputs include faces, medical imagery, surveillance footage and sound recordings.
… a processor subdivides a plurality of channels in the input feature data for the CNN into groups of intermediate channels and processes each individual, subdivided group of intermediate channels sequentially rather than processing all of the channels simultaneously.
Translation: Instead of handling all the data at once, the processor tackles smaller batches in a specific order.
Why smaller memory needs matter for phones
The filing says high memory needs get in the way of energy-efficient operation, especially on mobile devices and other hardware with limited power or space. Cutting how much data has to sit in memory, and how often the chip reaches out to main memory, targets that problem directly.
For you, success would look like face recognition or photo analysis running faster or on smaller, cheaper devices without a bigger memory chip. For AMD, which makes graphics and AI processors (the filing mentions GPUs and NPUs), it means its chips could handle more work with the same memory. The savings figures, about 33% of the storage and about 8% of the memory accesses, are the filing's own estimates, so treat them as targets rather than test results.
AMD's 47th filing we've tracked since May in the AI chip wars adds to earlier applications covering chips sharing memory and surviving a failed network card.
This filing has a short road to a real product. Splitting the work into groups is a change in how a program orders its steps, so it could run on chips that already exist. The buffer is the one piece that needs fast memory sitting close to the processor.
The savings numbers come from the document's own estimates, with no test results shown. The layer pattern it spells out (a small filter, a per-channel filter, another small filter) tells you the authors have a specific kind of image model in mind, so the gains depend on how well your model fits that pattern.
My read: modest, practical engineering. The shortest route to a product is a software update to the code that runs an image model, plus a chip with enough nearby memory to hold the running total. Nothing in the text calls for new hardware. That makes it routine in the best sense: useful, easy to adopt, and unlikely to excite anyone outside a chip lab.
Get our take in your Top Stories
Liked this breakdown? Add Patentlyze as a preferred source on Google, and our plain-English take shows up more often in your Top Stories the next time AMD patent news breaks.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
3 drawing sheets from US 2026/0310353 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in