Google Patents an AI Memory That Stops Growing as Text Gets Longer
Today's AI language models keep a note for every word they've seen, so their memory use only goes up. A Google patent application describes a model whose memory stays the same size however long it runs.
What Google's fixed-size AI memory actually does
Every time an AI writes its next word, it checks its notes on every word that came before. Those notes pile up. The longer the conversation or document, the more of a computer's fast memory the notes eat, which is a big reason long chats are costly to run and hard to squeeze onto a phone.
Google's patent application describes a different habit. Instead of adding a new note for every word, the AI keeps one fixed-size notebook and rewrites it as it goes, squeezing new information in and letting stale details fade. The notebook is the same size after ten words or ten thousand.
In the filing's own tests on four collections of text, this design scored better than competing designs. The filing also says the fixed size would suit phones and smart home devices that have little memory to spare.
Systems and methods for generating a network output using a neural network. In particular, the neural network includes one or more compressive attention layers that each “compress” information into a respective fixed size memory state.
Translation: The system uses special layers to shrink input text into a fixed memory size.
How the notebook gets rewritten word by word
Most of today's AI text models are built on attention, a trick where each new word is compared against everything written so far. To avoid redoing that work, the model stores a key-value cache (a running set of notes, one pair per earlier word). That cache grows with every word and has to sit in expensive, fast chip memory.
The claimed method swaps in a compressive attention layer. For each new word, the layer builds a query, a key, a value and a target vector (four learned summaries of the current word). It then runs a few optimization steps (small gradient-descent corrections, the same kind of math used to train AI) on a compression objective, which nudges a fixed-size memory so it better reproduces the target. The updated memory is then applied to the query to produce the layer's output.
The description and dependent claims add several details:
- A two-pass update: one pass folds in the key, the other folds in the value.
- A forget gate, a number that controls how much old memory is kept or wiped, for example when the topic changes.
- A chunked version that handles groups of words in parallel, so GPUs and TPUs (chips built for AI math) can run it efficiently.
Why a memory that stops growing matters for phones
If you've ever watched a chatbot slow down or hit a limit on a long document, memory is often part of the story. The filing says the growing cache can use up the fast memory on AI chips, which caps how long a text the model can handle and how many people a server can serve at once. A memory that stays one size could ease both limits.
It could also help AI move closer to your pocket. The description says a fixed size lets the model fit into phones, smart home devices and other gadgets with little memory, and lets engineers tune it to a specific device. That is the filing's claim, backed by its own charts. It is a patent application, not a shipping feature.
Google's 18th filing we've tracked since July in the AI chip wars adds to a run that includes one on predicting overheating and one on skipping blank video.
A notebook that stays the same size has to forget something. This design squeezes everything the AI has read into a fixed set of slots, so details from far back can blur or get overwritten. The forget gate in the filing is the dial that decides what goes. A full cache keeps every word's notes intact, and that perfect recall is what you give up.
There is a second cost. Each new word now triggers a small round of math to rewrite the notebook, in two passes, where the old approach simply filed a note away. The filing answers this by processing words in chunks so chips can work in parallel.
The filing's charts use perplexity (a score for how well a model predicts the next word) on four text collections, and they show lower scores than rival designs. They say nothing about how well the model digs out one exact fact buried deep in a long document. For phones and budget servers, the trade reads as worth it if that recall holds up. For jobs where a single missed detail matters, recall is what to test first.
Get our take in your Top Stories
Liked this breakdown? Add Patentlyze as a preferred source on Google, and our plain-English take shows up more often in your Top Stories the next time Google patent news breaks.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
4 drawing sheets from US 2026/0310664 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in