Microsoft Patents an AI That Spots Objects in Video You Can Teach Without Starting Over
Training an AI to recognize a specific object or situation normally takes weeks of labeled data and expensive model retraining. Microsoft's new patent describes a way to skip all of that, letting users teach the system on the fly with just a few example images.
What Microsoft's customizable object detector actually does
Imagine you run a factory and want your security camera's AI to flag workers who aren't wearing a hard hat. Today, teaching that AI the difference between "hard hat on" and "hard hat off" usually means hiring data scientists, collecting thousands of labeled photos, and waiting days or weeks for a new model to be trained. Microsoft's patent describes a system that sidesteps all of that.
The idea is to let you, the user, show the AI a handful of examples of exactly what you care about, even pulling them from a live video feed in real time. The system then figures out what to look for by comparing your examples against automatically generated "contrast" examples showing the opposite condition.
You don't need to be a machine-learning engineer. You point, you mark, and the detector learns. That applies to complex rules too, like spotting a scene where both a forklift and a pedestrian are present at the same time.
… comparing a first similarity measure between the first representation and the second representation to a second similarity measure between the first representation and the third representation; …
Translation: The system checks how closely an object matches the target against its opposite state.
How the contrast-state comparison identifies the right objects
The core of the patent is an object detection system built on an open vocabulary model (a type of AI that already understands a broad range of objects and concepts without being locked into a fixed list of categories). The twist is a customization layer on top of it.
When a user defines an object of interest and a specific state (for example, "person" + "wearing a mask"), the system automatically generates a contrast state ("person" + "not wearing a mask"). Both states are converted into embeddings, which are dense numerical fingerprints that capture the meaning of an image or concept in a form a computer can compare.
For any detection in a new image, the system generates its own embedding and then measures similarity to both the target-state embedding and the contrast-state embedding. Whichever is closer wins, and that determines the classification output. This comparison step is what lets the system distinguish nuanced conditions without needing a full model retrain.
The patent also covers:
- Handling logical combinations (AND/OR conditions involving multiple objects) by training a lightweight classifier on top
- Accepting user-supplied images as examples, rather than text descriptions alone
- Processing live video streams where users can mark examples in real time
- Using a generative model to synthesize contrast images when real examples are scarce
… supports customization of classifications without requiring the model to be retrained.
Translation: Users can teach the AI new visual rules on the fly without starting over from scratch.
What this means for businesses using AI-powered cameras
For any business that uses AI-powered cameras or automated image review, retraining a model every time the rules change is a real cost. A retailer adding a new product to watch for, a hospital changing its safety protocols, a logistics company onboarding a new type of package all of these currently require engineering time. A system that accepts a few marked photos as instructions collapses that cycle from weeks to minutes.
The live-video-marking capability is particularly interesting for security and industrial monitoring, where the conditions you care about are sometimes easier to point to than to describe. Whether Microsoft builds this into Azure AI services, Copilot-connected cameras, or some other product line, the underlying approach targets a real friction point in how organizations actually deploy computer vision today.
Microsoft's 26th filing in our AI vision coverage since May adds to a run that includes one spotting objects in video and one reading emoji to summarize meetings.
Every camera system watching a warehouse, retail shelf, or construction site eventually hits the same wall: the moment a business needs it to notice something specific, specialists have to spend months rebuilding it. That delay kills projects before they prove their value.
Microsoft's patent attacks that directly by letting a system learn from a handful of photos a regular user supplies, rather than from a massive technical overhaul. Teaching a system what "too crowded" or "wrong placement" looks like by showing it pictures is what determines whether a tool actually reaches the people who need it.
The honest question is whether this holds up when the distinction someone is trying to teach is subtle or inconsistent, because that is exactly where these systems fail in practice. The patent does not fully answer that. But the problem it targets is real, widespread, and costly enough that even a partial solution carries genuine weight.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
13 drawing sheets from US 2026/0301369 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in