Intel · Filed Mar 28, 2025 · Published Oct 1, 2026

Intel Patents a Chip That Shuts Off Its Own Processing Units During AI Tasks

Running an AI model on-device is a power hog, and most chips don't adjust mid-task. Intel's new patent describes a processor that fires up all its cores to get started, then hands off to a smaller cluster to finish the job.

A processor's internal architecture, including multiple cores, cache levels, and a power management circuit for AI tasks. Drawing from patent filing US 2026/0300178 A1.
A processor's internal architecture, including multiple cores, cache levels, and a power management circuit for AI tasks.
See all 40 drawings from this filing ↓
Publication number US 2026/0300178 A1
Applicant Intel Corporation
Filing date Mar 28, 2025
Publication date Oct 1, 2026
Inventors Sujit Mahto, Jayesh Gaur, Adithya Ranganathan, Shankar Balachandran, Sudhanshu Shukla, Sreenivas Subramoney, Eran Shifer
US classification 711/141
Status when we published On hold at the patent office (May 19, 2025)
Document 20 claims

How Intel's mode-switching chip handles AI word by word

When you ask an AI assistant a question, the chip inside your device has to generate an answer one word (or word fragment) at a time. That first word takes the most effort: the chip needs to look up a lot of information at once. But each word after that is a simpler operation, mostly reading from what it just figured out.

Intel's patent describes a chip that takes advantage of that difference. It uses all of its processor cores working together to produce that first word, then automatically switches to a smaller group of cores for every word that follows. The rest of the chip powers down or steps aside, burning far less energy.

The switch happens without you doing anything. A built-in power-management circuit watches what the AI model is doing and reconfigures the chip's internal connections on the fly. The goal is to get the same answer with less battery drain, which matters a lot for phones, laptops, and anything else running AI locally.

How the processor shifts cores and memory links mid-inference

The patent centers on a specific observation about how large language models (the AI systems behind chatbots and on-device assistants) actually work. Generating a response happens in two distinct phases.

The first phase, often called the prefill stage, is compute-heavy: the chip reads your entire prompt and sets up a compressed representation of it. This benefits from many processor cores working in parallel. The second phase, token generation (producing each word or fragment of the answer), is less compute-bound but reads a large amount of cached data repeatedly, which taxes memory bandwidth more than raw processing power.

Intel's design exploits this split by building a chip with two configurable operating modes:

  • Mode 1 (full-chip): All processor cores, the interconnect (the internal highway linking cores to memory), and all cache-coherence circuits (the hardware that keeps every core's local memory in sync) are active to produce the first output token.
  • Mode 2 (reduced-cluster): A subset of cores, a narrower interconnect path, and fewer coherence circuits take over for all subsequent tokens.

The switch is managed by a power management circuit baked into the chip itself. Critically, the cache-coherence circuits, hardware that normally prevents cores from reading stale data from each other, are partially disbanded when those extra cores go idle. Fewer active nodes means less coordination overhead and less wasted power keeping dormant hardware in sync.

What this means for AI on laptops and low-power devices

For you, the most direct payoff is battery life. Running an AI model locally, whether that's a writing assistant, a real-time translation tool, or a coding helper, currently drains a laptop or phone faster than almost any other task. A chip that automatically scales back to a leaner configuration after the hardest part of the job is done could meaningfully extend how long you can use those features unplugged.

Intel has been filing around on-device AI efficiency since at least 2024, and this patent fits that pattern. The design also has implications for data-center inference, where thousands of requests are processed simultaneously and energy costs are enormous. Shaving power from the token-generation phase, which is where most of the time in a long AI response is actually spent, is where real efficiency gains live.

Intel's 44th filing we've tracked in the AI chip wars since May adds to a run that includes faster memory lookups and blocking wasted data fetches.

Editorial take

If you run AI features on a phone or laptop, the device tends to get warm and the battery drains faster than it should. This patent describes a chip that automatically scales back its internal activity once the heavy lifting of starting an AI task is done, burning less power through the quieter work that follows.

The part worth understanding is that the savings come from winding down the coordination layer between processing cores, not just slowing the cores themselves. That coordination overhead is a real source of heat and battery drain, and most chips keep running it whether or not the moment requires it.

For a person using the device, the change shows up as a phone that stays comfortable in your hand during a long AI session and still has meaningful charge left by evening. You would never read about it in a spec sheet, but you would notice its absence on any chip that lacks it.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

40 drawing sheets from US 2026/0300178 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.