Microsoft · Filed Jun 11, 2026 · Published Oct 8, 2026

Microsoft Files Patent for Cutting the Memory Cost of AI Text Generation

AI tools that summarize and translate waste a surprising amount of memory by keeping duplicate copies of the same notes. A new Microsoft patent application shrinks one example from 27.4 GB to 8.65 GB, then down to 2.7 GB.

An AI system writes news headlines while sharing stored working memory across possible wording choices to save computing power. Drawing from patent filing US 2026/0310929 A1.
An AI system writes news headlines while sharing stored working memory across possible wording choices to save computing power.
See all 16 drawings from this filing ↓
Publication number US 2026/0310929 A1
Applicant Microsoft Technology Licensing, LLC
Filing date Jun 11, 2026
Publication date Oct 8, 2026
Inventors Yu YAN, Jiusheng CHEN, Ruofei ZHANG
US classification 704/257
Status when we published Waiting for an examiner (Jul 2, 2026)
Parent application is a Continuation of 17178385 (filed 2021-02-18)
Document 20 claims

What Microsoft's shared-memory text patent does

When an AI tool writes a summary or translates a sentence, it doesn't commit to one answer right away. It tries several draft versions at once and keeps the best ones, and today each draft carries its own copy of the same saved background notes. Those copies pile up fast. In the patent's own example, one model needs 27.4 GB of memory just for those notes.

Microsoft's filing describes sharing a single copy among all the drafts, which drops that figure to 8.65 GB in the same example. It also adds a memory watcher: when space gets tight, the system throws away some of the saved notes and recalculates them later if room opens up.

For you, the benefit would be indirect. The same computer chips could handle more requests at once, which is how AI text tools get cheaper and faster to run.

From the filing · CLAIM 1
… maintain, for the first layer, a number of beam-specific cache regions used during the beam search that is less than a beamwidth of the beam search based on the single layer-specific cache region that is shared.

Translation: The system saves memory by having multiple search paths share the exact same cached data instead of duplicating it.

How the decoder frees memory layer by layer

Text tools like summarizers and translators usually have two halves. An encoder reads the input and turns it into numbers, and a decoder writes the output one word at a time. To pick good wording, the decoder uses beam search (it keeps several of the most promising draft sentences alive at once, and the number of drafts is called the beamwidth, often 4 to 200). Each layer of the decoder also relies on a cache, meaning saved intermediate results so it doesn't have to recalculate them.

Normally every draft gets its own copy of that cache for every layer. The claims cover a cache-location table (a lookup list that tells each layer where its saved data sits in memory) so that all the drafts point to one shared region per layer. That leaves fewer draft-specific cache regions than the beamwidth. In the filing's example (batch of 128 inputs, 4 drafts, 1,024-word inputs, a BART model), the cache falls from 27.4 GB to 8.65 GB.

The second idea is a memory watcher. It checks the cache against a threshold, and if the cache grows too big, it:

  • Picks a lower layer's cache region and flags it as not needed in the table;
  • Lets new data overwrite that space;
  • Recomputes the dropped data later if memory frees up.

In the example this pushes the minimum cache to 2.7 GB, which leaves room for a much larger batch of inputs. The filing notes that recomputing can sometimes beat reading from the cache anyway, since each cache access takes two reads and one write.

From the filing · THE ABSTRACT
When an amount of memory utilized by one or more caches of the neural network model is determined to exceed a threshold memory size, a layer-specific portion of a cache associated with a layer of the neural network model is identified.

Translation: When memory usage gets too high, the system pinpoints specific cache segments tied to layers of the neural network.

Why cheaper memory means bigger AI batches

Memory limits how much work an AI chip can do at one time. If the saved notes take up less room, the same hardware can process bigger batches of requests, and that is how summarizing, translating and question answering get cheaper to run. The filing names those tasks, plus paraphrasing and question generation, as targets. You would never see this change directly, but you might notice it in speed or in what a service can afford to offer.

The approach also trades a little extra calculation for a lot of freed memory. The patent says the speed impact is minimal, though that comes from the company's own example rather than any independent test. It is a patent application, so it describes an idea, not a shipped product, and the filing does not tie it to a named Microsoft service.

Microsoft's 20th filing we've tracked since July in our AI chip wars watchlist follows one routing requests by answer length and one shifting storage by demand.

Editorial take

Memory is a bottleneck behind many AI text services. A chip can only hold so much at once, so every gigabyte spent on duplicate notes is a gigabyte that can't go toward serving more people. Cutting 27.4 GB to 8.65 GB in the patent's example, and down to 2.7 GB with the memory watcher, goes after that waste directly.

The fix matches the size of the problem. Sharing one copy instead of several is a simple idea, and the numbers the filing gives are large. The throw-away-and-recompute step costs extra calculation, and the filing says the speed hit is minimal, a claim nobody can check from paper.

Those figures come from one example setup, a single model with specific settings, so real savings would vary. A problem this basic and a cure this direct still make the filing matter most to the people who pay the server bills.

Get our take in your Top Stories

Liked this breakdown? Add Patentlyze as a preferred source on Google, and our plain-English take shows up more often in your Top Stories the next time Microsoft patent news breaks.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

16 drawing sheets from US 2026/0310929 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.