AMD Patents a Chip That Picks Its Own Best Code Path While Running
Every GPU instruction set contains dozens of ways to do the same job, and choosing the wrong one can waste a huge chunk of processing power. AMD is patenting hardware that auditions its own code paths on the fly and locks in the winner.
What AMD's runtime kernel-picker actually does for you
Every time a program hands a big math job to a GPU, the chip has to pick a specific recipe for doing that work. Those recipes are called kernels, and there are usually many versions of each one. Picking the wrong version means slower results, higher power use, or both, and the "right" choice can change depending on the job size, the data, and the hardware.
AMD's patent describes extra circuitry built into the chip itself that watches several kernel options run in real time, scores them on quality metrics like speed and efficiency, and then automatically switches to whichever one is winning. You don't have to do anything. The chip is essentially running a quiet competition among its own code recipes while your program is already executing.
The practical idea is that the chip stops relying on developers or drivers to pre-select the best approach. Instead, it samples performance during your actual workload and adapts on the spot.
sample, using profiling circuitry, a kernel space including multiple kernels to evaluate the multiple kernels based on one or more quality metrics; identify an optimal kernel from the multiple kernels in the kernel space; and use the optimal kernel for a computation task.
Translation: The chip tests multiple versions of code while running and instantly switches to the best performing one.
How the profiling circuitry samples and scores each kernel
The patent describes an accelerator (a specialized chip or chip block, like a GPU or AI processor) that contains dedicated profiling circuitry. During normal program execution, that circuitry doesn't just run a single kernel; it samples across a kernel space, meaning it evaluates multiple candidate implementations of the same computation.
Each kernel is judged against one or more quality metrics. The patent doesn't lock in specific metrics, but the language covers things like execution time, throughput, and resource usage. The profiling circuitry monitors and analyzes how each kernel behaves specifically during the live computation task, not in a synthetic pre-run benchmark.
Once the circuitry identifies the optimal kernel from that evaluated set, the accelerator commits to it for the task at hand. Key things to note about the design:
- All of this happens at runtime, not at compile time or driver installation.
- The selection is tied to the user's specific computation task, not a one-size-fits-all default.
- The profiling hardware is built into the accelerator itself, so it doesn't rely on external software to make the call.
This is different from existing approaches where software libraries (like AMD's ROCm or Nvidia's CUDA tuning tools) pre-select or offline-tune kernels before a workload runs. Here, the hardware is doing the tuning loop itself, in the moment.
The profiling circuitry monitors and analyzes how the multiple kernels behave during execution of the computation task.
Translation: Built-in monitors watch how different code options perform while the processor is actually doing work.
What self-tuning hardware means for GPU workloads
For AI training and scientific computing, kernel choice can swing performance by 20-40% on the same hardware. Today, getting that choice right requires software engineers to run offline benchmarks, update driver configs, or use auto-tuning libraries that add setup time. If AMD can move that decision into the chip itself, workloads could self-optimize without any developer intervention.
For everyday users, the impact would show up as faster AI inference on AMD hardware, better GPU utilization in games that use compute shaders, and lower power draw for the same output. Whether this circuitry appears in a discrete GPU, an AI accelerator card, or an integrated chip is an open question, but the patent covers all of those targets.
AMD's 19th filing we've tracked since June in our GPU rendering race watch adds to earlier work on finding triangles faster and stopping mid-scene stalls.
The patent describes dedicated circuitry built into a chip to automatically find the best way to run a computation during actual use, not beforehand. That means new silicon, not a software update, and new silicon has real costs in space and power that this document does not resolve.
Before this reaches a product, AMD would need to show that the efficiency gains justify those costs, and the company's existing software tools already do a version of this job. A hardware layer adds speed and flexibility, but it also means every piece of software that talks to the chip would need to understand and cooperate with it, which is its own large project.
The shortest path to a shipping feature probably runs through a lower-pressure chip design, like a combined processor and graphics unit, before it ever touches a standalone graphics card where every square millimeter of chip space is fiercely contested.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
7 drawing sheets from US 2026/0300012 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in