Sony Patent Uses Body Poses as Bookmarks to Search Motion Data
Imagine teaching a computer to recognize a dance move by simply striking a pose at the start and end, like pressing record and stop with your body. That's exactly the idea behind this Sony patent.
How Sony's pose-bookmarking motion search works
Picture trying to search a massive library of recorded movements, maybe for a video game character, a sports training app, or an AR experience. The problem is that your search recording is almost never perfectly clean. You shift your weight before you start, you hesitate after you finish, and all that extra motion clutters the search and throws off the results.
Sony's idea is to use specific body poses as bookmarks. You strike a designated "start" pose, perform the movement you want to search for, then strike an "end" pose. The system reads those two poses from sensors worn on your body and automatically clips out only the motion in between, ignoring the messy stuff before and after.
The result is a much cleaner snippet of motion data that can be matched against a database far more accurately. Think of it like using chapter markers in a video instead of scrubbing through the whole thing by hand.
How the system detects start and end poses to clip motion
The patent describes an information processing device built around a "motion information extraction unit." This component watches a continuous stream of body-movement data coming from motion sensors worn on multiple parts of the user's body, things like wrists, ankles, or a headset.
The extraction unit monitors that stream for two special identification poses:
- (p) Start identification pose, a specific body position the user strikes to signal "begin recording my search motion now."
- (q) End identification pose, a second specific position signaling "stop here."
The system also supports an optional preliminary notification pose, essentially a "get ready" signal that can precede either the start or end marker. Once the two bookmarks are detected, only the motion data captured between them is extracted and sent to the search engine. Everything outside that window is discarded.
That clipped segment is then compared against a stored motion database to find the closest matching recorded movements. Because the extra fidgeting and transitions are stripped away, the comparison is working with cleaner, more representative data.
What this means for motion capture and gesture interfaces
Motion search is only as good as the quality of what you feed it. Any system that tries to match human movement, whether for game animation libraries, sports analytics, physical therapy, or AR/VR interactions, suffers when the input is noisy. Sony's approach shifts the burden of cleanup from software algorithms to a simple human convention: strike a pose to start, strike a pose to stop.
For you as a user, this could mean a future where you search for a move in a fitness or game app the same way you'd use a voice command: with a deliberate, standardized signal your body already knows how to make. It's a small design idea, but in motion-based interfaces where precision matters, cleaner input data changes everything downstream.
This is a tidy, well-scoped idea rather than a sweeping platform shift. Sony is essentially solving a data-quality problem with a behavioral convention, let the user mark their own motion clips by posing. It's the kind of practical UX thinking that tends to end up in real products, particularly in PlayStation or Sony's XR hardware pipeline where motion input is already central.
Which company should we read for you?
We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.
Get one Big Tech patent every Sunday
Plain English, intelligent commentary, no hype. Free.
Editorial commentary on a publicly published patent application. Not legal advice.