Disney · Filed Mar 10, 2025 · Published Sep 10, 2026 · verified — real USPTO data

Disney Patents an AI System That Pins Masks to Objects Across Every Video Frame

Keeping a digital mask locked to a moving face or logo through thousands of video frames is one of the most tedious jobs in post-production. Disney has filed a patent for a system that learns to do it automatically, frame by frame, without manual touchups.

A sequence of images showing an object, like a face, being masked and tracked across different video frames. Drawing from patent filing US 2026/0268706 A1.
A sequence of images showing an object, like a face, being masked and tracked across different video frames.
See all 4 drawings from this filing ↓
Publication number US 2026/0268706 A1
Applicant Disney Enterprises, Inc.
Filing date Mar 10, 2025
Publication date Sep 10, 2026
Inventors Gaspard Zoss, Derek Edward Bradley
CPC classification 345/419
Grant likelihood Medium
Examiner LIU, GORDON G (Art Unit 2618)
Status Non Final Action Mailed (Aug 5, 2026)
Document 20 claims

What Disney's per-frame object masking actually does

A director films a scene and then, in editing, wants to blur an actor's tattoo, swap a brand logo, or add a digital costume overlay that sticks perfectly to the performer no matter how they move. Today, that kind of work often means someone on a visual-effects team painstakingly adjusting the mask on every single frame. It's slow, expensive, and easy to get wrong.

Disney's patent describes a system that watches the video, analyzes the object in each frame, and builds what's called a UV map: think of it as a coordinate grid painted directly onto the object's surface, like latitude and longitude lines on a globe. The system figures out where each part of that surface sits in every frame, so a mask drawn once on one coordinate point automatically follows the correct spot through the whole clip.

You tell the system which region of the object you want covered, and it handles the rest, producing a clean masked version of the full video sequence.

From the filing · CLAIM 1
… extract a plurality of visual features of the same object depicted in the image; and process the plurality of visual features to predict an image-space UV map of the same object depicted in the image; …

Translation: The system analyzes the object in every single frame to map its exact shape and orientation.

How the UV map ties each pixel to a surface point

The system takes a video sequence in which the same object (a face, a body, a prop) appears in every frame. For each frame, a neural network extracts visual features (patterns of color, shape, and texture that identify the object) and uses them to predict an image-space UV map.

A UV map is a standard computer-graphics tool: it assigns every visible pixel on an object a unique 2D coordinate, essentially flattening the object's surface into a grid. UV mapping is already used to wrap textures onto 3D models; this patent applies the same idea directly to 2D video frames without needing a 3D model at all.

Once the system has a UV map for each frame, an artist (or an automated input) specifies a region using those UV coordinates. A face-tattoo region, for example, is defined once by its position on the coordinate grid. The system then:

  • Looks up that UV region in each frame's predicted map
  • Translates it back to the actual pixel locations in that frame
  • Applies the mask to those pixels

The result is a masked video sequence in which the selected region is consistently covered regardless of motion, rotation, or lighting change, because the coordinates are tied to the object's surface, not to fixed pixel positions.

From the filing · THE ABSTRACT
… receive a masking input identifying a portion of the object to be overlaid by a mask, and mask, using the masking input and the image-space UV map predicted for each image, a respective portion of each image of the sequence of images corresponding to the identified portion to produce a masked sequence of images.

Translation: An editor points to what needs covering, and the software tracks that exact spot across the entire video.

What this means for VFX and broadcast production

Manual frame-by-frame masking is one of the most labor-intensive tasks in film and television post-production. Even with existing tracking tools, artists often spend hours correcting drift, especially when a subject turns, goes out of focus, or is partially hidden. A system that predicts surface coordinates from image features alone, with no pre-built 3D model required, could cut that time significantly.

For Disney specifically, this kind of technology has broad applications: protecting performer privacy during production reviews, enabling last-minute costume or branding changes after filming wraps, and feeding cleaner data into other AI visual-effects pipelines. Disney's filing activity in AI-assisted production tools suggests the company is building out a stack of automation for its studios rather than relying on third-party software.

This is the fifth Disney filing we've tracked in our AI vision coverage since July, adding to work like one on scene-by-scene show effects and one on spotting ride faults.

Editorial take

Frame-by-frame visual effects cleanup is one of the most expensive invisible costs in film and streaming production. When a digital effect needs to stay locked to a moving actor or object across hundreds of shots, small drifts accumulate, and fixing them manually means skilled artists billing significant hours on projects already under budget pressure.

Disney's patent targets that specific drain by having the system read raw image pixels and directly predict where each point on an object's surface sits, rather than rebuilding a full three-dimensional model first. Fewer steps in the process means fewer places where things go wrong mid-shot.

The credibility of this filing comes from how precisely it names the problem. Editors and compositors who spend their days on frame-level cleanup will recognize the pain immediately, and a solution sized to a real, chronic production cost carries more weight than one chasing a theoretical one.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

4 drawing sheets from US 2026/0268706 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.