Microsoft · Filed Jun 23, 2025 · Published Aug 6, 2026 · verified — real USPTO data

Patent Uses AI to Shrink Video File Sizes by Predicting Frame Changes

Compressing video without losing quality is one of the hardest problems in media tech. Microsoft's latest patent describes an AI system that pulls extra context from nearby frames to make that compression more efficient.

Microsoft Patent: Neural Network Video Compression Explained — figure from US 2026/0229019 A1
Figure from the official USPTO publication.
See all 11 drawings from this filing ↓
Publication number US 2026/0229019 A1
Applicant Microsoft Technology Licensing, LLC
Filing date Jun 23, 2025
Publication date Aug 6, 2026
Inventors Jiahao LI, Bin LI, Yan LU
CPC classification 706/12
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (May 7, 2026)
Parent application is a National Stage Entry of PCTUS2023084484 (filed 2023-12-18)
Document 20 claims

How Microsoft's AI video compression actually works

Imagine you're watching a video call and the picture gets blurry every time someone moves. That blur happens because video compression cuts corners on motion, it only looks at a single previous frame to guess what comes next, and sometimes it guesses wrong.

Microsoft's patent describes an AI system that works differently. Instead of just glancing at the most recent frame, it gathers richer contextual information by combining motion estimates, a reconstructed version of a reference frame, and additional learned features pulled from that reference. Think of it like a film editor who studies not just the last shot but also their notes and a rough cut before deciding what to keep.

The result is that the system can make better decisions about what parts of a new frame actually need to be stored versus what can be inferred, potentially shrinking file sizes or improving picture quality at the same bitrate.

From the filing · CLAIM 1
… determining, using a context extraction model, context information for the target frame based on the estimated motion information, the reference reconstructed frame, and the reference feature information; …

Translation: An artificial intelligence model analyzes previous video frames to figure out what is likely to change in the next frame.

How the context extraction model builds richer frame data

The patent describes a pipeline with two main neural-network models working in sequence.

First, a context extraction model takes three inputs:

  • Estimated motion information, a map of how pixels moved between frames (similar to what traditional video codecs call motion vectors)
  • A reference reconstructed frame, the previously decoded frame stored for comparison
  • Reference feature information, learned internal representations (think of these as the AI's own notes about the reference frame, captured during a prior encoding step)

By combining all three, the context model builds a richer picture of what the current frame should look like before encoding even begins.

Second, a frame coding model uses that context to either compress the current frame into a bitstream or reconstruct it during playback. The key claim is that feeding feature-level information from the reference frame, not just the pixel image, gives the coder more signal to work with and reduces the bits needed.

The underlying approach is sometimes called learned video compression, a research area where neural networks replace or augment traditional hand-coded compression rules like those in H.264 or H.265.

From the filing · THE ABSTRACT
In this way, richer context information can be extracted for coding and thus the coding efficiency can be improved.

Translation: This method gathers more detailed data about the video so it can be compressed into a much smaller file size.

What this means for streaming and video call quality

Traditional video codecs like H.264 and HEVC are heavily engineered by hand and have been refined over decades. Learned or neural codecs are a younger field, and one of their persistent challenges is making full use of the information already available from neighboring frames. This patent targets that gap directly by injecting richer, feature-level reference data into the compression loop.

For you as a viewer, the promise is sharper video at the same internet speed, or the same quality at lower data cost. For Microsoft, which operates Azure video services, Teams, and Xbox cloud gaming, more efficient video coding has direct cost and quality implications across a huge user base.

Editorial take

This is a solid entry in the competitive learned-video-compression space, not a dramatic departure from existing research directions. The core idea of feeding learned feature representations into a coding loop is well-established in academic literature, so the novelty here likely lies in the specific model architecture and training setup rather than a conceptual breakthrough. Worth tracking if you follow video codec standardization, but not a headline moment on its own.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

11 drawing sheets from US 2026/0229019 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.