Nvidia · Filed Mar 10, 2025 · Published Sep 10, 2026 · verified — real USPTO data

Nvidia Patents an AI System That Writes Scripts and Voices the Characters Automatically

Nvidia has filed a patent for a system that takes a simple description, writes a full script, assigns AI agents to each character, and outputs a finished voiced conversation, all without a human in the loop.

An AI system generates scripts and character voices, showing the flow from input data to final audio output. Drawing from patent filing US 2026/0268890 A1.
An AI system generates scripts and character voices, showing the flow from input data to final audio output.
See all 18 drawings from this filing ↓
Publication number US 2026/0268890 A1
Applicant NVIDIA Corporation
Filing date Mar 10, 2025
Publication date Sep 10, 2026
Inventors Francesco Ciannella, Davide Marco Onofrio, Jose Rafael Valle Gomes da Costa
CPC classification 704/260
Grant likelihood Medium
Examiner THOMAS-HOMESCU, ANNE L (Art Unit 2656)
Status Docketed New Case - Ready for Examination (May 16, 2025)
Document 20 claims

How Nvidia's AI turns a prompt into a voiced conversation

A podcaster sits down to record, but instead of guests, the entire conversation is generated by AI from a one-line prompt. That is the kind of workflow Nvidia is building toward with this patent.

You give the system some basic details: how many characters you want, what roles they play, maybe a topic or a short description of the scene. From there, a chain of AI agents takes over. One agent writes the script, another acts as a director coordinating the performance, and individual character agents each generate speech for their lines. Voice cloning is used so each character sounds distinct, natural, and emotional.

The result is a complete audio conversation produced end-to-end by AI. No voice actors, no editors, no separate steps you have to manage yourself.

From the filing · CLAIM 1
… generating, using one or more language models and based at least on the information, a script associated with the conversation; generating, using the one or more language models and based at least on a first scene of the script, first audio data representative of first synthetic speech for the first character …

Translation: The system uses AI to write a script and create realistic spoken voices for the characters.

How the screenwriter, director, and character agents divide the work

The system described in the patent works through a layered hierarchy of AI agents, each with a specific job.

Step one: writing the script. A "screenwriter agent" takes your input (topic, character count, roles, descriptions) and produces a structured script broken into scenes. This is handled by a large language model (an AI trained on text, similar to the kind that powers chatbots).

Step two: directing the performance. A "director agent" reads that script and coordinates how each scene is performed. It decides which character speaks when and passes scene-level instructions to the individual character agents.

Step three: generating speech. Each character has its own AI agent responsible for producing synthetic speech for that character's lines. The agents use voice cloning (a technique that reproduces specific vocal qualities, tone, and emotion) to make each voice sound distinct and human rather than robotic.

The final output stitches together the audio from all character agents into a coherent conversation. The patent covers both the generation of the script and the generation of the audio, treating them as a single connected pipeline rather than separate tools.

From the filing · THE ABSTRACT
… where the character agents may use voice cloning to cause the voices of the characters to sound realistic, human, and/or emotional.

Translation: AI voice cloning makes the generated character voices sound like real, emotional humans.

What this means for AI-generated audio and virtual assistants

For anyone building AI assistants, training simulations, or automated content pipelines, this kind of end-to-end system removes a significant amount of manual work. Right now, producing a multi-character voiced conversation typically requires separate tools for writing, voice synthesis, and audio assembly. Collapsing that into a single prompt-to-audio pipeline is a meaningful reduction in effort.

For everyday users, you might encounter this in AI tutoring apps, interactive games, or podcast-style content where the "hosts" are fully synthetic. The voice-cloning component is the part most people will notice first: whether the characters sound convincingly human or slip into the uncanny valley will determine whether this technology feels useful or unsettling.

That makes this Nvidia's eighth filing we've tracked since July on our watchlist for AI models working in teams, after one covering a local task-routing hub and one covering a self-fixing code writer.

Editorial take

The concrete payoff here is time. Producing a multi-character audio conversation today involves a writer, a voice director, voice actors or a text-to-speech setup, and then someone to edit it all together. This patent describes collapsing that chain into a single input.

Whether you notice the difference depends entirely on voice quality. The system leans on voice cloning to carry the emotional weight, and that technology is still inconsistent enough that a bad output breaks the illusion immediately. The architecture of agents-writing-for-agents is interesting, but the listening experience is what will determine whether anyone actually uses this.

Nvidia's steady investment in agentic AI shows up across a range of filings lately, and this one fits that pattern. Audio generation is a less-covered corner of that bet, and this patent plants a clear flag in it.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

18 drawing sheets from US 2026/0268890 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.