Adobe Patents a Text-to-Animation System That Moves Characters Like Their Bodies Actually Would
Type a description like 'a heavyset man jogs across the room' and Adobe's system doesn't just generate a jog -- it generates a jog that looks right for that body. Most animation AI ignores body shape entirely; this one bakes it in from the start.
How Adobe's body-shape motion AI actually works
A character animator types a sentence describing how a character should move. The software generates the motion automatically, but the result looks generic because the system doesn't know whether the character is tall, short, broad, or slight.
Adobe's patent describes a way to fix that. You give the system two things: a description of the motion you want and a description of the body doing it. The system figures out both the shape of the body and how that body would naturally move at the same time, rather than doing them separately. The result is animation that fits the character, not just the action.
For creators, that means less cleanup work. Right now, auto-generated motion often has to be manually adjusted because a walk cycle designed for a stick-figure skeleton looks wrong on a stocky character. This approach tries to get it right the first time.
… jointly predicting, by the processing device, a shape parameter and a plurality of motion tokens inferred by a machine learning system based on the description …
Translation: The software simultaneously calculates how a character's body type and their intended movement should interact.
Inside Adobe's joint shape-and-motion prediction pipeline
The system takes a text prompt covering both a movement (say, 'a person waves goodbye') and a body shape (say, 'tall and slender') and feeds that into a machine learning model trained to handle both at once.
Instead of generating motion and then trying to fit it to a body afterward, the model jointly predicts two things simultaneously: a shape parameter (a compact numeric description of body proportions) and a set of motion tokens (discrete chunks representing segments of the movement, similar to how language models break text into word-pieces).
Those motion tokens go through a step called finite scalar de-quantization (essentially: converting compressed motion codes back into detailed movement data). The shape parameter is then projected into the same mathematical space as the motion data and merged with it, creating what the patent calls shape conditioned motion features. Think of it as stamping the body's proportions directly onto every frame of movement.
Finally, the model decodes those combined features into actual motion data, the kind that can drive a 3D character rig, with the body shape influencing how the limbs travel, how weight shifts, and how posture adjusts throughout the sequence.
The shape conditioned motion features are decoded by the machine learning system into the motion data, which models the motion with specific variations caused by the shape.
Translation: The system creates animation data that adjusts how a character moves based on their unique physical build.
What this means for animators and AI-generated characters
For anyone using AI tools to generate character animation, the practical gap today is that auto-generated motion tends to look like it was made for a default, average body. Fixing that mismatch is tedious manual work inside tools like After Effects or Mixamo. A system that bakes body shape into the generation step would mean fewer hours of cleanup for the animator.
Adobe's product lineup already includes motion-related tools inside Premiere Pro, After Effects, and its Firefly AI family, so the research fits squarely into an ongoing push to automate production tasks that currently eat up time. Big Tech patent news in the AI animation space shows a pattern of companies racing to close exactly this gap between 'text to motion' and 'text to believable motion for a specific character,' and this filing is Adobe's entry into that race.
That makes this Adobe's 18th filing we've tracked since May in our controllable AI image work, following filings like separate prompt compositing and text-driven 3D edits.
The concrete payoff here is time. An animator who uses a text-to-motion tool today gets a starting point, but they spend real hours adjusting it to match the character's body. If Adobe ships something close to what this patent describes, that adjustment step shrinks or disappears.
The harder question is whether the system holds up on unusual or extreme body shapes, where training data is typically thin. A motion model that works well for near-average proportions but falls apart on outliers would still leave animators doing cleanup, just on a narrower set of characters. The patent doesn't address that limitation directly.
Still, the architecture choice to predict shape and motion at the same time rather than in sequence is the right instinct. Sequential approaches accumulate error at each step; joint prediction gives the model a chance to let each constraint inform the other from the start.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
9 drawing sheets from US 2026/0253298 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →