Nvidia Patents an AI Plug-In That Translates Video Calls Into Sign Language Avatars
Nvidia has filed a patent for a plug-in that would sit inside any video conferencing platform and automatically translate spoken conversation into sign language, delivered through an animated avatar on screen. No special hardware required.
How Nvidia's sign language video-call plug-in would work
Ever tried to follow a fast-paced video meeting when captions are delayed, wrong, or simply missing? For deaf and hard-of-hearing people, that frustration is the default experience, not the exception. Nvidia's new patent describes a plug-in designed to fix that.
The idea is straightforward: audio or text from a video call gets fed through an AI translation system, which converts it into the user's preferred sign language. A virtual avatar then appears on screen and performs that sign language in real time, so a deaf participant can follow the conversation visually without relying on error-prone auto-captions.
Crucially, the patent describes this as a plug-in for existing platforms, meaning it would theoretically work inside tools like Zoom or Teams rather than requiring everyone to switch to a new app. Your sign language preference is stored and applied automatically for every call.
… generate first control data to control animation of at least a portion of a first avatar to perform signing corresponding to the sign language translation data …
Translation: The system creates digital instructions that tell a virtual character how to move its hands to sign the spoken conversation.
How the avatar animation and translation modules connect
The patent describes a software plug-in that can be loaded into a video conferencing client application (think Zoom, Google Meet, or Microsoft Teams). Once active, it spins up one or more sign language translation modules for the duration of a call.
Those modules take incoming communication channel data (which can be spoken audio, text, or even an incoming sign language video stream) and translate it into a target sign language based on the individual user's stored preference. The patent specifically notes the system can handle translation from spoken language or from one sign language to another, covering multilingual signing scenarios.
The translation output drives avatar animation control data: a set of instructions that tells the rendering engine exactly how to move a virtual character's hands, arms, and face to perform accurate signing. The plug-in then takes over a portion of the video conferencing user interface to display that avatar alongside the normal call view.
Key components described in the claim:
- Processing circuitry that instantiates translation modules per session
- A sign language preference setting tied to the individual user
- Control data generated to animate an avatar's signing in real time
- Direct integration with the client app's UI layer to display the result
… translate incoming communication channel data from a first language mode (e.g., spoken language or sign language) to a target language mode corresponding to a sign language preference for the user.
Translation: The software converts audio or video input into the specific sign language format that the user prefers to see.
What this means for deaf and hard-of-hearing video callers
For the roughly 70 million deaf people worldwide who use sign language as a primary language, automated captions are often the only accessibility option in video calls, and those captions frequently stumble on accents, technical terms, or fast speech. A real-time signing avatar would be a more natural and reliable alternative for that audience, and because the patent describes a plug-in architecture, it could theoretically be added to existing platforms without rebuilding them from scratch.
Nvidia's position here is notable: the company's GPU hardware already powers most of the AI inference that makes real-time translation and 3D avatar rendering practical, so a software layer on top of that infrastructure is a logical extension of what they already sell to cloud and enterprise customers. Accessibility-focused AI filings are appearing more frequently among new Big Tech patents, and this one puts Nvidia squarely in a space that has so far been dominated by smaller startups and academic research groups.
Claim 1 is written broadly enough to cover any processor-based system that translates call audio into sign language and shows an avatar doing the signing inside a video platform's UI. That scope is wide: granted with minimal narrowing, it could create friction for any video conferencing company or accessibility startup building a similar feature without a license from Nvidia. The claim skips any restriction to a particular sign language, platform, or avatar type, making it a territorial stake in the whole category of accessible communication tooling. Whether the patent office finds enough prior art to cut that scope down will matter a great deal for who else can compete here.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
11 drawing sheets from US 2026/0237323 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →