Nvidia Patents Software That Fills In Video Frames It Never Saw
Every video is really just a series of still images played fast enough to look like motion. Nvidia is patenting a system that uses AI to invent the frames between those stills, making footage smoother or filling gaps without ever touching a camera.
How Nvidia's system invents frames that weren't filmed
Imagine you're watching a slow-motion replay, but the original clip wasn't filmed in slow motion. To stretch it out, someone would have to create extra frames that were never actually recorded. Today that's hard to do without obvious blurring or stuttering.
Nvidia's patent describes an AI system that learns how things move, particularly how bodies and objects shift position over time, and uses that knowledge to predict what the missing frames should look like. The result is a second, smoother version of the original video with extra frames inserted in the right places.
The system is built around neural networks trained specifically on motion patterns, so it isn't just blending adjacent frames together. It's making an educated guess about where everything was headed.
How pose-trained networks infer the missing motion
The patent centers on a processor whose arithmetic logic units (the chips that do the actual math) are set up to run one or more neural networks capable of frame inference, meaning they calculate what a video frame should contain even if it was never captured.
The key detail is how those networks are trained: using temporal pose representations, which are structured descriptions of how a body or object is positioned at different points in time. Think of it like a connect-the-dots sketch of a person's skeleton across multiple moments, captured from real video. The network studies thousands of these motion sequences and learns the rules of how things move.
At inference time (when you actually feed it a video), the network takes the frames it does have, figures out the motion trajectory, and generates plausible new frames to sit between them. The output is a second video with a higher frame count than the original.
- Input: an existing video clip
- Processing: neural network infers motion using learned pose data
- Output: a new video with additional synthesized frames inserted
What this means for video upscaling and real-time AI
Frame interpolation already exists in consumer TVs and video editing software, but it's notoriously bad at handling fast motion, complex backgrounds, or human bodies in action. By grounding the AI in pose-based motion data, Nvidia's approach is targeting exactly the cases where older methods fall apart.
Nvidia makes the GPUs that run most of the world's video processing and AI workloads. If this technique makes it into DLSS (Nvidia's existing frame-generation technology for games) or into video production tools, it could meaningfully raise the quality ceiling for AI-generated slow motion, video upscaling, and real-time frame boosting in everything from gaming to streaming.
Frame generation is already one of Nvidia's biggest selling points for its RTX GPU line, and this patent suggests the company is investing in making that AI smarter about human and object motion specifically. That's a real gap worth closing. The abstract is sparse, but the pose-training detail is the meaningful part.
Which company should we read for you?
We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.
Get one Big Tech patent every Sunday
Plain English, intelligent commentary, no hype. Free.
Editorial commentary on a publicly published patent application. Not legal advice.