Nvidia · Filed Jun 9, 2026 · Published Oct 1, 2026

Nvidia Patents an API That Lets Programs Control How a GPU Splits Its Work

GPUs do their best work when thousands of tiny tasks run in parallel, but today's software has limited say over exactly how that traffic gets routed. Nvidia's new patent describes an API that hands that control back to the programmer.

A compute unit's grid of thread blocks is organized into block clusters, each containing multiple thread blocks. Drawing from patent filing US 2026/0300006 A1.
A compute unit's grid of thread blocks is organized into block clusters, each containing multiple thread blocks.
See all 73 drawings from this filing ↓
Publication number US 2026/0300006 A1
Applicant NVIDIA Corporation
Filing date Jun 9, 2026
Publication date Oct 1, 2026
Inventors Ze Long, Kyrylo Perelygin, Harold Carter Edwards, Gokul Ramaswamy Hirisave Chandra Shekhara, Jaydeep Marathe, Ronny Meir Krashinsky, Girish Bhaskarrao Bharambe
US classification 719/328
Status when we published Waiting for an examiner (Sep 2, 2026)
Parent application is a Continuation of 17955070 (filed 2022-09-28)
Document 20 claims

What Nvidia's thread-cluster scheduling API actually does

A modern GPU is basically a city with thousands of tiny workers. Right now, when a program sends work to the GPU, the GPU's own internal scheduler decides which worker handles which task, and the program mostly just waits. That arrangement is fast in general, but it can be inefficient for programs that know exactly how their work should be divided.

Nvidia's patent describes a programming interface (a formal way for software to talk to hardware) that lets a program specify the scheduling policy before the GPU starts working. Instead of the GPU guessing, your code can say: split these thread clusters (groups of parallel tasks) across the processing units this way, not however you feel like it.

The practical effect is that software developers writing GPU-intensive applications, think AI training or scientific simulations, get a new lever to pull. They can tune how work is distributed to match their specific problem, rather than hoping the GPU's defaults are a good fit.

From the filing · CLAIM 1
cause one or more block clusters, each block cluster comprising a plurality of blocks of threads, to be performed concurrently by a plurality of compute units of an accelerator, wherein performance of the blocks of threads is distributed between the plurality of compute units according to a scheduling policy indicated to the API.

Translation: Software can now tell the graphics chip how to break up and run tasks at the same time.

How the API routes thread blocks across compute units

The patent centers on a single idea: an API call (a standardized instruction a program sends to the hardware) that includes a scheduling policy as one of its parameters.

Here's the structure the patent describes:

  • Block clusters: the GPU's work is organized into clusters of "blocks," where each block is itself a bundle of parallel threads (individual instructions running simultaneously). Think of a block as a team and a cluster as a division of teams.
  • Compute units: the physical processing sections of the GPU that actually execute those threads. A modern GPU can have hundreds of them.
  • Scheduling policy: the rule the GPU follows when deciding which compute unit handles which cluster. This is what the new API exposes. Instead of that rule being hidden inside the GPU driver, the calling program declares it explicitly.

The patent is written around CUDA, Nvidia's programming platform for GPU computing (the standard toolkit most AI and scientific software uses). So this isn't a hypothetical future interface; it's designed to slot into an ecosystem that millions of developers already work in.

The core claim is that running these clusters concurrently (at the same time) across multiple compute units, with a programmer-specified distribution rule, produces outcomes that the GPU's own default scheduler might not achieve.

From the filing · THE ABSTRACT
Apparatuses, systems, and techniques to execute CUDA programs. In at least one embodiment, an application programming interface is performed to determine a scheduling policy of one or more blocks of one or more threads.

Translation: The patent describes new ways to run Nvidia programming code using custom scheduling rules.

What this means for GPU-heavy AI and HPC workloads

For most consumer software, this patent changes nothing you'd notice day-to-day. The GPU in your laptop will keep running games and video edits just fine without any of this. The audience here is the relatively small group of engineers writing the low-level code that underpins AI model training, data-center inference, and high-performance computing, exactly the markets where Nvidia makes most of its money.

For those developers, gaining explicit control over scheduling policy is meaningful. Poorly distributed work across compute units causes some units to sit idle while others are overloaded, wasting expensive hardware time. A scheduling API lets a program avoid that waste for specific, well-understood workloads. Whether Nvidia exposes this in a way that's easy to use, or buries it in a niche CUDA extension few people touch, will determine how much real-world impact it has.

Nvidia's 49th filing we've tracked since July in the AI chip wars watchlist follows earlier work on a compression math fix and formatting data in the connector, adding another piece to how it handles AI data.

Editorial take

Giving programmers manual control over how a chip organizes its own work sounds like pure upside, but the cost is real: when a human makes a worse call than the chip would have made on its own, the result is slower and more expensive, and the chip gives almost no feedback explaining why.

For the narrow group this targets, people writing custom software for machines that cost tens of thousands of dollars, that risk is probably acceptable. They have strong reasons to chase every efficiency gain, and they have the skill to use a sharp tool carefully.

The design trades simplicity for headroom, which is a reasonable bet for specialists and a bad one for everyone else. Whether that bet pays off depends almost entirely on how clearly Nvidia explains the tool, because a dial with poor instructions tends to cause more damage than having no dial at all.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

73 drawing sheets from US 2026/0300006 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.