Nvidia's Patents For Text-to-3D Characters, and Where They Point
This watchlist tracks Nvidia's filings on turning text prompts, photos, and video into rigged, animatable 3D characters and full digital humans, replacing manual modeling and studio capture. Together, the filings point to steady research into generating simulation-ready characters from whatever input is easiest to get.
21 filings
· tracking since May 2026 · latest Sep 2026 · updates weekly
based on all tracked filings in this watchlist · refreshes every week
Nvidia is filing patents around turning simple inputs like text, photos, and video into full 3D characters that can move, emote, and be used in games or virtual spaces.
The filings cluster most heavily around two areas: building realistic 3D human bodies from everyday photos or text descriptions, and making those characters move in ways that look natural and respond to instructions.
What’s new in Nvidia's text-to-3D characters
a dated entry each week this watchlist moves · older entries stay archived
Sep 17, 2026 1 filing joined
This week's filing focuses on making digital faces move and speak in real time by reading audio. The push is toward characters that react to sound as it happens, not after the fact.
This week's filing covers a tool that reads a script and automatically draws out a full set of scenes. The focus shifts from building 3D characters to planning and picturing whole stories visually.
This week's new filing focuses on cleaning up raw 3D scans so they can actually be used as digital models. Nvidia is working on ways to take messy scan data and turn it into something a computer can work with properly.
The filing pace inside Nvidia's text-to-3D characters
The focus areas inside Nvidia's text-to-3D characters
the problems Nvidia keeps filing on · each with its three newest filings · new filings join every week
Text and Photos Into 3D Characters 6 filings
Building a usable 3D character from words or a simple photo is slow and requires expert tools. These filings cover automated pipelines that take text descriptions or everyday photos and produce 3D characters ready to move and be placed in simulations.
Making a virtual character move in a way that looks natural is hard to do without hand-crafting every action. These filings cover AI approaches that generate and predict character motion from simple instructions or learned rules about how bodies move.
Creating realistic digital humans normally requires expensive studio capture sessions. These filings cover systems that build full-body digital human models from ordinary 2D photos or video footage of real people.
Flat images and video hide information about shape, depth, and body structure that is needed to make characters move realistically. These filings cover systems that recover or build a hidden 3D model from images, video, or generated pictures.
After establishing how to generate 3D characters from text and video, this filing targets the animation layer: getting facial movements, mouth, eyes, teeth, to sync with speech in real time rather than requiring manual keyframing or post-processing.
Automated storyboarding from scripts removes a manual bottleneck before 3D character generation even starts, letting creators move directly from story to asset production.
Turning a raw 3D scan of a real object into a clean, usable digital model is one of the messiest problems in computer graphics. Nvidia's new patent describes a machine-learning approach that may make that process significantly more automatic.
Real-time sign language output would let deaf participants follow conversations at natural speaking pace rather than waiting for captions. This filing shows Nvidia pursuing accessibility as a direct application of its character animation tech.
Generating animations on demand from high-level commands sidesteps the need to pre-record or manually craft every possible movement a 3D character might perform, letting game engines request poses and transitions dynamically instead.
Training a single model on both motion reconstruction and generation lets the network learn shared patterns that improve performance on each task, rather than building separate specialized tools for partial data versus text-driven invention.
The pipeline shifts from static character generation to autonomous movement: after learning from motion-capture data, the AI must produce convincing motion without direct video reference, relying only on reward signals for realism.
Constraint-based motion generation lets artists specify only key behavioral rules, what must happen and what must not, while the system synthesizes natural motion between keyframes, cutting manual pose adjustment work.
The pipeline now includes a correction layer that catches anatomical implausibility before output, using learned body proportions rather than manual fixes.
A 3D character generated this way would actually look like the reference photo instead of drifting into the AI's generic defaults. The patent reveals how Nvidia fuses text and image inputs so neither one overwhelms the other.
After establishing how to generate 3D characters from text and reference images, this filing shows how to choreograph multiple characters together using motion from video references, controlling both identity and movement in a single pass.
Extracting emotion intensity from speech audio lets the system sync facial expressions to vocal tone rather than words alone, closing the gap between what a character says and how it actually sounds.
The text-to-3D pipeline needs generated images to stay editable after creation. This filing confirms Nvidia is embedding 3D structure directly into the generation process rather than trying to reconstruct it afterward.
A single foundation model trained on 2D internet photos eliminates the need for separate specialized networks for face, hands, and body, compressing what normally requires studio equipment into one unified system.
Ready-to-animate characters need consistent body geometry across poses. This filing shows how to extract that consistency from messy, uncontrolled photos instead of requiring studio capture rigs.
The 3D character pipeline so far relies on controlled capture data; this filing shows how to work backward from random internet photos with conflicting angles and lighting to reconstruct consistent human geometry.
Generating clothing that behaves correctly under physics simulation requires predicting how fabric deforms and moves. This filing shows how Nvidia chains text-to-3D generation with physics-aware training to produce garments that don't need manual rigging.
Multi-angle video capture lets the system build skeletal structure and surface geometry in a single pass, collapsing what normally requires separate motion-capture and modeling stages into one automated workflow.
Animators would skip the manual rigging step entirely, with the diffusion model generating skeletons and joint controls automatically alongside geometry. This shrinks the pipeline from separate modeling, rigging, and animation phases into one model output.
Within the character pipeline, this filing shows how Nvidia routes text descriptions directly to physics-compatible geometry, cutting the traditional bottleneck of manual asset creation that slows down simulation environments.
Rigging animated characters requires skeleton structures and joint definitions that most AI models generate poorly. This filing shows how to extract usable skeletal data directly from the diffusion model's output, bypassing manual rigging work.
Questions readers ask
What problem is Nvidia trying to solve with these patents?
Across the filings, Nvidia is trying to replace manual 3D modeling and rigging with an AI pipeline that starts from a text prompt, a photo, or video footage and outputs a character ready to animate and simulate. The filings describe research direction, not a shipped product, so the exact workflow may change before anything reaches users.
Is this about generating game characters or something else?
The filings describe simulation-ready 3D characters and full digital humans generally, without naming a specific game or product. The same pipeline could apply to games, film, or other 3D content, but the patents themselves stay focused on the underlying generation method rather than a named use case.
How is this different from existing 3D character tools?
Existing tools generally need a trained artist to build and rig a character by hand, or a controlled capture studio to scan a real person. Nvidia's filings describe skipping both steps, generating a rigged, animatable character directly from a text description, ordinary video, or uncontrolled 2D photos instead.
Will these patents turn into real Nvidia products?
Patents describe what a company is exploring, not what it has committed to ship. This watchlist shows Nvidia filing multiple related pipelines around text-to-3D and photo-to-3D characters, which signals sustained research interest, but there is no guarantee any of it becomes an announced product.
Want this weekly breakdown for a company we don't cover?
Patentlyze Pro →
The weekly email: the best of Big Tech's filings, in plain English. Free.