Microsoft Patents AI That Captions Your Body Movements During 3D Video Calls
Microsoft is working on a system that watches how you move during a 3D video call and automatically generates live captions describing your motion, the same way subtitles describe speech. It could turn a videoconference into something closer to a real-time movement analysis session.
What Microsoft's motion-captioning system actually does
A physical therapist watches a patient do a shoulder exercise over a 3D video call. She can see the video, but there is no easy way to log what the patient actually did, how far the arm extended, or whether the posture was correct. That gap is what this Microsoft patent tries to close.
The system scans a 3D video feed of a person, figures out where their joints are in each frame, and feeds that data to an AI model. The AI reads the pattern and writes a plain-text caption describing what the person is doing. That caption floats next to the person's 3D image in real time, like subtitles for movement.
Physical metrics, things like range of motion or speed, could also be calculated and shown on screen. The system is designed for 3D video conferences, not standard flat video, so it has enough depth information to track your body position accurately.
… receiving, from the machine learning model, a description of a motion of the subject; and displaying the description of the motion of the subject with a rendering of one of the plurality of 3 D representations of the subject.
Translation: Getting the motion description from AI and showing it on screen next to the 3D video.
How joint coordinates become live text descriptions
The system starts with a 4D mesh (a 3D model of a person's body that updates frame by frame over time) captured during a videoconference. From that mesh, it extracts the positions of the person's joints at each point in time.
Those joint positions are converted into text-based representations, essentially structured descriptions of where each body part is in space, angle and all. The patent calls these "parameterized models," meaning the body is described as a set of numbers rather than pixels.
A machine learning model then receives those numbered descriptions alongside a prompt context (background instructions that tell the AI what to look for, such as "this is an exercise session" or "watch for these specific movements"). The AI produces a natural-language caption, for example "right arm raised to 90 degrees" or "forward lunge, incomplete extension."
The caption is overlaid directly next to the person's 3D image in the conference view. Separately, the system can also compute physical metrics such as speed or range of motion and display them in real time alongside the video feed.
… a caption describing the subject’s motion is overlaid proximal to the 3 D representation of the subject.
Translation: Text explaining how someone is moving appears right next to their 3D video feed.
What this means for remote physical therapy and coaching
For remote physical therapy, sports coaching, or rehabilitation, the gap between "I can see you" and "I can measure what you're doing" has always been a problem. This system would let a therapist or trainer read a live text log of a patient's movement without manually annotating anything, which could make remote sessions far more useful.
There is also an accessibility angle: captioning motion the way we caption speech could help people who rely on text-based descriptions follow along in a 3D video environment. Whether this ends up inside a future version of Microsoft Teams or a specialized health platform, the underlying idea applies any time precise body movement matters in a remote setting.
Microsoft's 11th filing in the AI simulation filings we cover since May follows work like one mapping farm cause and effect and rebuilding 3D bodies from photos.
The design trade here is real: converting joint positions into text before handing them to a language model is a clever shortcut, but it also throws away a lot of information. The raw 3D mesh carries continuous spatial data; once you flatten it into text tokens, some precision is gone, and the AI's caption is only as good as how well those text descriptions capture the original motion. For rough-and-ready descriptions of common exercises, that is probably fine. For clinical-grade measurement, it probably is not.
The prompt-context mechanism, where you tell the AI what kind of session this is, adds flexibility but also adds a failure point. A prompt tuned for physical therapy could misread a dance rehearsal; a generic prompt might not be specific enough for either. That context has to come from somewhere, and the patent is quiet on how it is set or who sets it.
Still, the core idea of treating body movement as something worth captioning, the same way speech has been captioned for decades, is a practical one. several Microsoft filings on 3D conferencing and spatial computing this year suggest this is part of a broader push, and motion captioning is at least a coherent piece of it. The tradeoffs are real but not fatal.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
5 drawing sheets from US 2026/0301321 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in