Qualcomm Files Patent for Chatbot Chips That Skip the Work They Don't Need
Every time an AI chatbot adds a word to its reply, its chip has a little more to remember. Qualcomm's patent application describes a way to reserve that space once and only do math on the part that's filled in.
What Qualcomm's skip-the-blank-space trick does
You ask your phone's AI assistant a question, and it starts typing out an answer one small chunk of text at a time. With every chunk, the chip has to look back over everything said so far, and that pile of notes keeps growing.
That growth is a headache for the software that prepares AI programs to run, because it works best when every size is fixed in advance. Qualcomm's filing gets around it by reserving a notebook big enough for the longest conversation the model can handle. It then adds a simple mask, a marker that says which pages are actually filled in.
The chip does math only on those pages and skips the blank ones. The filing also describes keeping the notes on the AI chip itself, so they don't have to be shipped back and forth with the phone's main processor after every chunk.
define a maximally sized static data structure and a variable data mask, the variable data mask indicating a valid section of the maximally sized static data structure …
Translation: The chip sets up a large container for data and uses a mask to highlight only the parts it actually needs.
How the mask marks which memory counts
A chatbot writes its reply one token (a chunk of text averaging about four characters) at a time, running the same calculation again for each one. To avoid redoing old work, it keeps a running memory called the KV cache (a table of notes about every token so far), and that table gets one row longer with each step. A static compiler (a tool that plans and schedules chip work before the program ever runs) prefers data whose size never changes, so a growing table is a poor fit.
The claims cover a two-part fix. First, define a maximally sized static data structure: a table fixed at the longest context the model can handle (the description says that ranges from two thousand to several hundred thousand tokens). Second, define a variable data mask, a marker for which part is currently filled in. The description says the mask can be as simple as one number, the count of valid rows, which starts at the prompt length and goes up by one per token. The chip then computes using only that valid section, and the compiler can skip the math and memory reads for the rest.
The filing also describes these add-ons:
- Keeping the KV cache on the AI accelerator (a chip built for this kind of math) instead of copying it to the host device, such as a phone's main processor, and back each round.
- Suppressing those copies (the hardware transfer jobs known as DMA, or direct memory access) after the first one, with compiler settings linking the skipped input to the earlier output.
- Generating code ahead of runtime that does only the minimal computation on the valid data.
A processor-implemented method for adapting autoregressive inference of large language models (LLMs) for static compilation …
Translation: This method helps artificial intelligence models run more efficiently on fixed hardware setups.
Why on-phone chatbots could answer sooner
If you use an AI assistant on a phone, what you feel is waiting. The filing says its techniques may improve processing efficiency and reduce latency (the delay before and between words appearing). Skipping math on empty memory is one way to cut that delay. Not shuttling a growing table back and forth between chips, which the filing says can leave the accelerator sitting idle, is another.
The filing names a smartphone as one example of the host device, though it also says the host and the accelerator can be separate machines. Its first claim is worded broadly, covering a mask applied to a fixed-size data structure, with the chatbot details in the later claims. The filing gives no speed figures, so how much faster things get is unknown.
Qualcomm's 58th filing we've tracked in the AI chip wars since July adds to a run that includes storing giant models as small edits and phones reporting their AI limits.
You will never see this on a spec sheet. If it works as described, you would notice it as a chatbot that starts answering sooner and keeps its pace in a long conversation, because the longer the chat, the more notes the chip has to handle.
The idea is simple: reserve the space once, mark what's in use, and don't waste effort on the blank part. Simple ideas that save work on every single word a chatbot writes can add up quickly for anyone talking to an AI on a phone.
The filing offers no speed numbers, so the size of the benefit is a guess. Treat it as behind-the-scenes engineering that could make your assistant feel quicker, and judge it by whether replies actually arrive faster.
Get our take in your Top Stories
Liked this breakdown? Add Patentlyze as a preferred source on Google, and our plain-English take shows up more often in your Top Stories the next time Qualcomm patent news breaks.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
8 drawing sheets from US 2026/0310941 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in