Microsoft Patents a Chain-of-Networks Approach to Spotting Objects in Video
Spotting a small, fast-moving object in a blurry video clip is one of the harder problems in computer vision. Microsoft's latest patent applies a relay-race approach, where each piece of an AI model hands off what it learned to the next, building up a richer picture of what's happening across time.
How Microsoft's video AI tracks blurry, tiny, or hidden objects
Imagine you're watching security camera footage of a parking lot at night. A small figure moves quickly across the corner of the frame, blurred by speed and half-hidden by a parked car. A single still-image AI would likely miss it entirely. That's the problem Microsoft is trying to fix here.
The patent describes an AI model that's broken into a series of smaller sub-models, each one responsible for a different video frame. As one sub-model finishes analyzing its frame, it passes a summary of what it found to the next sub-model, which uses that context when examining the following frame. By the end of the chain, the model has a much more complete picture of what's moving, where, and how fast.
The result can be used for tasks like labeling what's in a video, drawing boxes around objects, or segmenting the scene into regions. Think of it as giving the AI a short-term memory so it can connect the dots across time, rather than treating each frame as a brand-new puzzle.
… assigning individual video frames of the video to respective subnetworks of a machine learning model; extracting initial features from respective video frames of the video; inputting the initial features to respective assigned subnetworks of the machine learning model; propagating intermediate features among the subnetworks …
Translation: The system splits video frames across multiple smaller networks and passes data between them.
How the subnetworks pass context from frame to frame
The system works by distributing individual video frames across a series of subnetworks (smaller, specialized components inside a larger AI model). Each subnetwork handles one frame at a time, but it doesn't work in isolation.
Here's the flow the patent describes:
- The system pulls a sequence of video frames from a clip.
- Each frame is assigned to its own subnetwork in a fixed order.
- The subnetwork extracts initial features from its frame (essentially a compressed mathematical description of what it sees: edges, shapes, colors, motion patterns).
- Those features, plus intermediate features passed forward from the previous subnetwork, are combined and processed.
- The final subnetwork in the chain produces final features that represent the whole sequence.
- Those final features feed into a task-specific output layer for detection, classification, or segmentation.
The key innovation is the propagation of temporal features, meaning information about earlier frames is carried forward through the chain rather than discarded. This is what helps the model handle motion blur (where a fast-moving object appears smeared), occlusion (where an object is partially hidden), and small object size (where a single-frame snapshot may not carry enough pixels to be reliable).
The patent doesn't tie itself to a specific model architecture, which keeps the claims fairly broad.
Objects can be difficult to detect in videos due to motion blur, occlusion, and/or having a relatively small size.
Translation: Video analysis struggles when objects are blurry, hidden, or too small.
What this means for security cameras and video AI tools
Video AI is everywhere: in security systems, sports analytics, autonomous vehicles, and content moderation. The persistent weak point in most deployed systems is exactly what this patent addresses. Small, fast, or partially hidden objects are exactly what cause false negatives in real-world footage, and those misses carry real consequences in safety or surveillance contexts.
For everyday users, this kind of technology could mean a home security camera that actually catches the small detail that matters, or a video conferencing tool that correctly tracks multiple people moving around a room. The patent is filed broadly enough that if granted, it could cover a wide range of implementations where temporal context is passed between AI subnetworks during video analysis, which describes a pattern common to many modern video AI products.
Microsoft's 33rd filing we've tracked since May in AI teams working together builds on ideas like meeting role assistants and a self-managing security trio.
Claim 1 covers any computer-implemented method that takes a video, splits it into frames, sends those frames through a chain of processing stages, and passes context forward from each stage to the next before producing a final output. That structure is the backbone of virtually any modern video analysis system, making the claim remarkably wide.
In practice, a granted patent on that design could reach any competing product built around the same relay logic, where earlier processing informs later processing across a sequence of frames. Surveillance cameras, self-driving car systems, and streaming platforms all depend on exactly this approach to catch fast-moving or partially hidden objects.
Whether the claim holds up depends on what earlier work the patent office finds, but the scope Microsoft is asking for is large enough that the answer will matter well beyond this one product.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
17 drawing sheets from US 2026/0301374 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in