Google Patents a Step-by-Step Noise-Removal System for Generating AI Video
Google has filed a patent for a method of building video from scratch using a process that starts with pure noise and slowly carves out recognizable footage, one pass at a time. It's the same broad family of techniques behind today's AI image generators, extended to moving pictures.
How Google's video diffusion process actually builds footage
Today's AI video tools struggle because generating moving images is far harder than generating a single still: every frame has to look right on its own and flow naturally from the frames around it. Google wants to change that with a system that works the way a sculptor might, starting from a rough block and cutting away a little bit of material on each pass until the final shape appears.
The system takes your text prompt or other input, then spins up a rough, noisy placeholder for the whole video. A model works through that placeholder again and again, each time making the result a little less random and a little more like real footage, until the output video is done.
The result, if it works as described, is a video generation pipeline that can be conditioned on almost any input, whether that's a description, an image, or a reference clip, and that refines its output incrementally rather than trying to get everything right in a single step.
… generating an output video by updating the current intermediate representation at each of a plurality of iterations, wherein the updating comprises, at each iteration: processing an intermediate input for the iteration comprising the current intermediate representation using a diffusion model that is configured to process the intermediate input to generate a noise output …
Translation: The system creates video by repeatedly refining a digital image through a process that removes noise step by step.
How the noise iterations refine each video frame
The patent describes a system built around diffusion models, a class of AI that learns by training on data that has been progressively scrambled with random noise, then learning to reverse that scrambling. For images, this process is now well understood. For video, the challenge is that the model has to handle time as a dimension alongside height and width.
Here's how the pipeline works:
- Input reception: The system accepts a conditioning input, for example a text description or a reference image, that defines what the output video should look like.
- Initialization: A starting representation is created. Think of it as a canvas filled with static, the way an untuned TV screen looks.
- Iterative refinement: At each of many iterations, a diffusion model reads the current noisy representation and predicts what noise to subtract. The system then updates the representation using that prediction. Each pass brings the video closer to something coherent.
- Output: After enough iterations, the accumulated updates yield a complete video.
The key engineering choice is doing this work iteratively, meaning the model never has to produce a perfect result in one shot. Instead, small, compounding corrections accumulate into a finished video, which mirrors how image diffusion models like Stable Diffusion or DALL-E 3 operate.
What this means for Google's AI video ambitions
Diffusion-based image generation went from a research curiosity to a product feature inside two years. If the same trajectory holds for video, a reliable diffusion pipeline for footage could feed directly into tools like Google's Veo video generator or future versions of Search that synthesize video answers on the fly. For you as a user, it means AI-generated video clips that are more consistent frame-to-frame and more controllable from a simple prompt.
The filing is relatively early-stage and the first independent claim was canceled, which signals the patent is still being shaped through examination. Still, it represents Google staking out territory in the AI video generation space at the foundational method level, and it sits alongside a broader wave of latest Big Tech patents in AI-generated media that are defining who owns the underlying techniques for this technology.
That makes this Google's 42nd filing we've tracked since May in the AI photo editing race, following their work on erasing people from scenes and pulling concepts from photos.
Running a video through dozens of refinement passes costs real money in server time and electricity, and those costs multiply quickly because video files are far heavier than still images. Google's bet is that the quality improvement justifies that expense, but the efficiency gains the business needs may not arrive on the schedule the bills do.
The deeper vulnerability here is legal. A rejected foundational claim in this filing suggests an examiner found earlier work covering similar ground, which means Google will likely need to narrow what it is actually claiming before this patent grants. A narrower patent means narrower protection, and a design this expensive to run deserves strong legal cover to make the investment worthwhile.
Whether the tradeoff reads as worth it depends almost entirely on how fast the underlying hardware gets cheaper. The approach is sound in principle, but it is priced for a future that has not arrived yet.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
6 drawing sheets from US 2026/0253399 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →