Samsung · Filed Apr 10, 2025 · Published Jul 16, 2026 · verified — real USPTO data

Samsung Patents a System That Turns Your Voice Into Full-Body Avatar Movements in Real Time

What if your 3D avatar could wave its hands, tilt its head, and shift its weight just by listening to you talk? Samsung has filed a patent for exactly that, a system that generates full-body avatar gestures from your voice alone, running entirely on your device.

Samsung Patent: Real-Time Audio-Driven 3D Avatar Gestures — figure from US 2026/0203983 A1
Figure from the official USPTO publication.
Publication number US 2026/0203983 A1
Applicant Samsung Electronics Co., Ltd.
Filing date Apr 10, 2025
Publication date Jul 16, 2026
Inventors Byeonghee Yu, Danke Xie, Srinivasa Reddy Algubelli, Siva Penke
CPC classification 345/473
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Apr 23, 2025)
Parent application Claims priority from a provisional application 63745729 (filed 2025-01-15)
Document 20 claims

How Samsung's voice-to-avatar animation actually works

Imagine you're on a video call but instead of your face on camera, you're represented by a 3D character. Right now, most avatar systems either freeze in place or loop generic animations. Samsung's patent describes a way to make that avatar actually move the way a real person would while talking, nodding, gesturing, shifting posture, all driven by your voice.

The system listens to your audio in small chunks, extracts patterns from your speech (things like rhythm, emphasis, and tone), and feeds those into a model that predicts how a body would naturally move while saying those words. It then generates animation frames that are played in sync with your actual audio, so the avatar's movements match when you speak.

All of this runs on your device, not on a remote server. That means your voice data doesn't have to travel anywhere, and the whole process can happen fast enough to feel live.

How the on-device model predicts joints from speech

The patent describes a pipeline that processes incoming audio in real time, split into short sequential chunks stored in audio buffers (think of them as brief recording windows that refresh continuously).

For each chunk, an audio encoder (a model trained to read speech signals) pulls out speech features, not the words themselves, but acoustic properties like cadence, stress, and intonation. These features are handed off to a gesture generation model, which has been trained to predict two things:

  • The 3D position of the body's center (where the torso is in space)
  • The orientation of every body joint (how the head, neck, shoulders, elbows, wrists, and so on are angled)

From those predictions, the system generates animation keyframes (the specific poses an animation engine uses to smoothly draw motion between points in time). Those keyframes are stored in memory and played back in sync with the original audio so lip movement, head nods, and hand gestures all line up with what the speaker is actually saying.

Critically, the claim specifies an on-device model, meaning the inference runs locally. This keeps latency low and avoids sending raw audio to the cloud.

What this means for video calls and virtual presence

For video calls, virtual meetings, or social apps built around avatars, the difference between a talking head that loops the same idle animation and one that actually moves naturally is the difference between something you tolerate and something that feels like real presence. Your avatar looking alive without needing a camera pointed at your body is a meaningful quality-of-life shift for anyone who prefers not to be on camera.

Running the model on-device rather than in the cloud also matters for privacy. Your voice is the input, but it never has to leave your phone or headset. For Samsung, this fits a broader pattern of pushing AI processing onto its own hardware, which keeps users inside its ecosystem and reduces dependence on external services.

Editorial take

This is a genuinely practical patent, not a moonshot. Audio-driven avatar animation is a real problem in video calling and social apps, and doing it on-device is the right architectural call. The interesting question is whether Samsung ships this in Galaxy AI features or something like its video call apps, the patent is specific enough that it reads like something already in development, not just a defensive filing.

Which company should we read for you?

We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.

Get one Big Tech patent every Sunday

Plain English, intelligent commentary, no hype. Free.

Source. Full patent text and figures from the official USPTO publication PDF.

Editorial commentary on a publicly published patent application. Not legal advice.