Microsoft Patents an AI That Learns Tasks by Watching Someone Else Do Them
Microsoft wants to turn a single recorded walkthrough into an on-demand AI tutor, so the next person tackling the same task can ask questions and get answers pulled directly from that footage.
How Microsoft's demo-video AI would guide you through a task
Right now, if you need to learn a new procedure at work, you usually rely on a written manual, a colleague who happens to be free, or a training video you have to scrub through on your own. None of those options can answer a follow-up question mid-task.
Microsoft's patent describes a system where someone records themselves doing a task once, and an AI breaks that video into meaningful chunks, labels each chunk, and stores everything in a database. When you try the same task later, the system can pull the most relevant clip or caption and use it to answer your specific question in real time.
Think of it as a self-building FAQ for any procedure someone has bothered to record. The person who wrote the original "how-to" doesn't have to do any extra work, and you get a guide that can actually talk back.
… segmenting the demonstration video into multiple video segments based at least on the one or more contextual signals; generating augmentation data associated with individual video segments of the demonstration video using a multi-modal generative model; and storing the augmentation data in a database …
Translation: The system breaks down the demonstration video into smaller parts and uses AI to create helpful data stored for later use.
How the system segments, captions, and retrieves the right clip
The patent describes a pipeline that starts when a first user records a demonstration video of themselves completing a task. The system then processes that video using what the patent calls contextual signals (cues like speech, on-screen text, or detected actions) to split the footage into logical segments, each representing a meaningful step.
For each segment, a multi-modal generative model (an AI that can handle both images and text at the same time, similar to how GPT-4o processes pictures alongside words) generates augmentation data: things like keyframes (the most useful still images from that segment) and natural-language captions describing what's happening.
All of that gets stored in a database. Later, when a second user attempts the same task and asks a question, the system performs a retrieval step: it searches the database for the augmentation data most relevant to that query, then feeds it into a generative model to produce a direct, contextual answer.
The net effect is a loop:
- Record once
- Auto-index with AI
- Retrieve on demand
- Answer in plain language
The second user never has to watch the whole video; they just ask what they need.
When another user attempts to perform the task, selected augmentation data can be retrieved and used to prompt a generative model to answer user queries relating to the task.
Translation: When someone else tries to do the same task, the saved data helps an AI answer their questions along the way.
What this means for AI-powered workplace guidance tools
The practical problem here is real: institutional knowledge constantly walks out the door when experienced employees leave, and written documentation rarely keeps pace with how procedures actually evolve. A system that can convert a screen recording or workshop video into a queryable knowledge base could save organizations significant time on training and troubleshooting, without requiring anyone to author structured content from scratch.
For Microsoft, this fits neatly into its existing push to embed AI into productivity tools. The approach also touches on a broader shift in how AI assistants are being designed, moving away from generic chatbots toward systems grounded in specific, verified demonstrations. That shift is showing up across many new Big Tech patents in the AI-assisted workflow space, and Microsoft's filing adds another concrete implementation to that growing stack.
The problem this addresses is genuine and underappreciated: most organizations are terrible at capturing procedural knowledge in a form that's actually searchable, and the cost of that failure shows up daily in re-training, errors, and time spent asking the same questions over and over. The approach here is proportionate to the problem. Grounding AI answers in a specific recorded demonstration rather than a general language model makes the system much less likely to hallucinate a step that doesn't exist in the actual procedure. Whether the technical pipeline described can handle the messy, poorly lit, audio-noisy videos that real workplaces produce is the real test.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
17 drawing sheets from US 2026/0244470 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →