Samsung · Filed May 26, 2026 · Published Sep 24, 2026 · verified — real USPTO data

Samsung Patents a Video Repair System That Uses Distance Sensing to Fill Missing Frames

Removing an unwanted object from a video is still one of the messiest problems in editing. Samsung's new patent tries to fix the flicker and inconsistency that plagues today's automated tools by feeding depth information into the restoration process.

A camera captures two slightly different views of a scene, demonstrating the input for generating missing frames. Drawing from patent filing US 2026/0289754 A1.
A camera captures two slightly different views of a scene, demonstrating the input for generating missing frames.
See all 16 drawings from this filing ↓
Publication number US 2026/0289754 A1
Applicant Samsung Electronics Co., Ltd.
Filing date May 26, 2026
Publication date Sep 24, 2026
Inventors Sungho LEE, Jaewoong SOH, Yelin HAN, Youngjin YOON, Yeoul LEE, Narae CHOI, Woongil CHOI
CPC classification 382/254
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Jul 14, 2026)
Parent application is a Continuation of PCTKR2024018854 (filed 2024-11-26)
Document 20 claims

What Samsung's depth-guided video restoration actually does

When a video editor removes something from footage, whether a camera rig, a wire, or an unwanted person in the background, the software has to invent what should be behind it, for every single frame. Current tools often produce results that flicker or look inconsistent because each frame gets treated somewhat separately.

Samsung's approach adds depth information, a rough map of how far each part of the scene is from the camera, to help the system understand the scene in three dimensions. That depth data guides how each restored frame relates to the frame that came just before it, keeping the fill-in visually consistent over time.

The result is supposed to be video where removed objects leave behind a patch that holds together from one moment to the next, rather than swimming or shifting in ways that look artificial. Think of it as giving the repair tool a sense of the room, not just the wallpaper.

From the filing · CLAIM 1
extracting an input feature including an image feature and a depth feature, based on a first frame to be reconstructed, a mask, and a depth frame; …

Translation: The system gathers visual details and depth information from the damaged video and its depth map.

How depth features guide the frame-by-frame fill process

The patent describes a step-by-step process a device runs when it needs to reconstruct a damaged or masked region of a video frame.

  • Feature extraction: For the frame being repaired (the "first frame"), the system pulls out two kinds of information: an image feature (what the pixels look like) and a depth feature (how close or far each part of the scene is from the camera, derived from a depth frame, which is basically a grayscale distance map).
  • Reference frame selection: The system picks a prior frame (the "second frame") that is related to the first. Because video is continuous, the frame that just played contains strong hints about what the missing region should look like.
  • Depth-guided alignment: The features of the current frame are aligned, or warped, to match the prior frame's output, but the alignment is steered by depth data from both frames. This means the system accounts for three-dimensional position when deciding how to map one frame onto another, not just flat pixel similarity.
  • Feature modification: The image feature is then updated using the depth feature of the aligned input, so the visual fill-in is shaped by spatial context.
  • Reconstruction: The corrected feature set generates the final restored frame.

The depth signal acts as a spatial anchor. Instead of guessing what belongs in the masked region based only on surrounding pixels, the model knows roughly where in three-dimensional space that region sits, which helps it pull consistent information across frames.

From the filing · THE ABSTRACT
… aligning the input feature of the first frame with an output feature of the second frame, based on depth features of the first frame and the second frame, …

Translation: It matches the broken frame with a previous frame by comparing how far objects are in both scenes.

What this means for video editing and AI restoration tools

Video inpainting, filling in erased or missing parts of footage automatically, is one of the most sought-after tools in film post-production, surveillance cleanup, and consumer video apps. The core problem is temporal consistency: frame-by-frame AI fills tend to drift or flicker because the model treats each frame in isolation. Depth-guided alignment is a practical attempt to close that gap by giving the model spatial memory.

For everyday users, this kind of tech is what eventually ends up in phone camera apps and desktop editors that let you erase objects with a tap. Samsung keeps filing on on-device AI video processing, which suggests this depth-based approach could be aimed at future Galaxy devices that handle restoration locally, without uploading footage to a cloud server. Whether this specific method makes it to a shipping product is impossible to say from the patent alone.

Samsung's 24th filing we've tracked since July in the AI photo editing race follows one that reshapes its own selection and one borrowing from your other photos.

Editorial take

The problem this patent targets is real and expensive. Professional VFX studios spend significant time and money on clean plates (background footage without unwanted objects) precisely because automated removal tools still struggle with temporal consistency. Every time the fill-in flickers, a human has to fix it by hand.

Depth information is a sensible lever to pull here. If the system knows that the area being repaired sits three feet behind a foreground subject, it can make better-informed guesses about what texture and lighting should appear there, and it can do that consistently across frames. The logic is sound.

That said, the quality of the depth frame is the weak link the patent does not resolve. Consumer-grade depth sensing, from phone cameras or monocular depth estimation models, is noisy. A patched process is only as good as the map it relies on, so real-world performance could vary widely depending on what feeds the depth input.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

16 drawing sheets from US 2026/0289754 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.