Qualcomm · Filed Apr 7, 2025 · Published Oct 8, 2026

Qualcomm Files Patent for Storing Giant AI Models as Small Edits to a Shared Copy

Giant AI models keep many near-identical expert copies sitting in memory. Qualcomm's patent keeps one shared copy and stores each expert as a small set of differences.

A large AI program split into one common master copy and small lists of changes, cutting the storage needed on a device. Drawing from patent filing US 2026/0311068 A1.
A large AI program split into one common master copy and small lists of changes, cutting the storage needed on a device.
See all 17 drawings from this filing ↓
Publication number US 2026/0311068 A1
Applicant QUALCOMM Incorporated
Filing date Apr 7, 2025
Publication date Oct 8, 2026
Inventors Mahdi NAZM BOJNORDI
US classification 706/10
Status when we published Waiting for an examiner (May 7, 2025)
Document 20 claims

What Qualcomm's shared-copy AI patent does

Ever kept ten nearly identical copies of the same document, each with a few edits, and wondered why you didn't just save one and note the changes? Many large AI models have a similar habit. They are built from groups of expert sub-models that share the same shape and differ only in their stored numbers, and every copy takes up room in memory.

Qualcomm's new patent describes keeping one shared set of numbers and, for each expert, storing only the small differences. In the filing's own example, a stored value of 4 becomes a shared 7 plus a difference of minus 3. Small differences take fewer bits to store, so there is less data to shuffle around.

The conversion happens after the model is trained, so the model would not need retraining. The filing says the pieces get added back together at the processor right before the math happens.

From the filing · CLAIM 1
… determine, based on a difference between each expert weight of the respective plurality of expert weights associated with each expert model and the plurality of base values of the shared base, a respective delta map for each expert model …

Translation: It figures out the exact math differences needed to turn the shared foundation into each specific expert model.

How the shared base and delta maps work

The patent targets Mixture of Experts models, AI systems that split their work across many specialist sub-networks and send each piece of input (called a token) to only a few of them. Every expert has the same structure. Only their weights, the learned numbers inside, differ.

The claimed core has three steps:

  • Pick a shared base: one set of base values worked out from the experts' weights.
  • For each expert, subtract the base from its weights to get a delta map, a grid of small differences.
  • Combine base and deltas to produce each expert's base-delta representation.

The filing also describes optional extras. Experts can be sorted into groups with k-means clustering (a method that piles similar items together) so each group shares a base. Oversized differences can be clamped to a cap, such as negative 2 to 2. The system can also list candidate setups, each defined by how many bases to keep and how many bits each difference gets, then pick the one with the lowest mean absolute error (the average size of the mistakes introduced). In the example, a setup with three bases and 2-bit differences scored best, at about 0.41.

The tradeoff is some extra arithmetic to rebuild the weights in exchange for less memory traffic.

From the filing · THE ABSTRACT
… generate, based on combining the plurality of base values and the respective plurality of delta values associated with each expert model, a respective base-delta representation.

Translation: The system then builds the final model version by putting the shared base together with its unique adjustment map.

Why smaller experts could help phone AI

Memory is a major bottleneck for running big AI models. The filing notes that expert models can outgrow a device's main memory, and that jumping between experts strains the data pipe between memory and processor. Storing experts as small differences means less data to move, which is why the filing pitches the idea for phones, laptops and servers alike.

For you, the payoff could be larger AI features running on more modest hardware, if the idea works as described. The filing says no retraining is needed, which would make it easier to apply to models that already exist. The costs are extra work to rebuild weights on the fly and small errors from clamping, so accuracy results on real models would decide whether it holds up.

Qualcomm's 57th filing we've tracked in our AI chip wars watchlist since July adds to a pattern that includes one on phones reporting AI limits and one on partial model updates.

Editorial take

Among the ideas in this filing, this one sits close to the shippable end. It is a software-style recipe: take a finished model, run a one-time conversion, and add the pieces back together at the processor. The filing lists ordinary chips (CPUs, GPUs and dedicated AI processors), and no new hardware is described.

What has to exist first is a model built from many experts, plus proof that squeezing it does not hurt answer quality. The only worked example here is a toy table of eight experts with 32 numbers each, and the error figures come from that toy. Nothing in the text shows results on a real chatbot.

The shortest route to a product is a conversion step in the software that prepares a trained model for a device. That is a small step compared with designing a new chip, which makes the idea practical to try.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

17 drawing sheets from US 2026/0311068 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.