New Google Patents · Filed Mar 31, 2025 · Published Oct 1, 2026 · verified — real USPTO data

Google Patents a System That Keeps Track of Who's Talking in Long Audio Recordings

Every AI transcription service has the same embarrassing problem: in a long meeting recording, the system might call the same person 'Speaker 1' in one section and 'Speaker 2' in another. Google's new patent tackles exactly that.

A system processes audio from multiple speakers into a diarized transcript, identifying who said what in a conversation. Drawing from patent filing US 2026/0300624 A1.
A system processes audio from multiple speakers into a diarized transcript, identifying who said what in a conversation.
See all 6 drawings from this filing ↓
Publication number US 2026/0300624 A1
Applicant Google LLC
Filing date Mar 31, 2025
Publication date Oct 1, 2026
Inventors Quan Wang, Neeraj Gaur, Guanlong Zhao, Parisa Haghani
CPC classification 704/9
Grant likelihood Medium
Examiner ISLAM, MOHAMMAD K (Art Unit 2653)
Status Docketed New Case - Ready for Examination (May 2, 2025)
Document 20 claims

How Google's speaker-tracking AI handles long conversations

Every time an AI tries to transcribe a long meeting or podcast, it faces a messy problem: the recording is too long to analyze all at once, so the system chops it into chunks and processes each one separately. In one chunk, your boss might be labeled 'Speaker A.' In the next chunk, the AI starts fresh and accidentally calls her 'Speaker B.' The transcript is technically accurate word-by-word, but the speaker labels are a jumble.

Google's patent describes a system that fixes this after the fact. A neural network (a type of AI modeled loosely on how the brain connects information) looks at all the separately processed chunks and figures out which labels should actually match across the whole recording. When two chunks disagree about who is who, the system swaps the labels in one chunk to make everything consistent.

The result is a single, clean transcript where the same person carries the same label from start to finish, even in a recording that runs for hours.

From the filing · CLAIM 1
obtaining a plurality of audio data segments characterizing a conversation between two or more speakers; for each respective audio data segment, generating, using a speaker diarization model, a corresponding short-form diarized transcript …

Translation: The system breaks a long conversation audio file into smaller chunks and labels who is speaking in each piece.

How the neural network reconciles mismatched speaker labels

The patent describes a two-stage pipeline for handling long audio recordings with multiple speakers.

Stage one: chunked diarization. A speaker diarization model (diarization means splitting audio by speaker identity) processes each audio chunk independently and produces a short-form transcript. Each transcript contains the words spoken and speaker tokens, which are placeholder labels like 'Speaker 1' or 'Speaker 2' that indicate who said what within that chunk.

Stage two: reconciliation. Here is where the core invention lives. Because each chunk is processed independently, the label assignments are arbitrary and inconsistent across chunks. A sequence processing neural network (a type of AI that reads data in order, similar to how a language model reads sentences) then examines the full set of chunk transcripts and resolves the conflicts. It does this by permuting speaker tokens, meaning it systematically swaps label assignments in specific chunks so that the same real-world person ends up with the same label everywhere.

The patent specifies that the network generates a reconciled long-form diarized transcript as its output. The design avoids re-processing the raw audio, working instead purely from the text-and-label outputs of the first stage, which keeps the computation manageable for arbitrarily long recordings.

From the filing · THE ABSTRACT
… generating a reconciled long-form diarized transcript based on the corresponding short-form diarized transcript generated for each respective audio data segment.

Translation: An artificial intelligence network stitches the smaller labeled chunks back together into one unified transcript.

What this means for transcription and voice AI products

Automatic transcription tools are already widely used for meetings, legal depositions, medical consultations, and podcast production. The speaker-label inconsistency problem is one of the main reasons these transcripts still require human cleanup, especially for recordings longer than about 30 minutes. A reliable fix would make AI-generated transcripts far more useful as standalone documents.

For Google, this capability fits naturally into products like Google Meet, Google Workspace's transcription features, and any voice AI that processes multi-speaker audio. Google keeps filing around long-form audio intelligence suggests the company sees speaker understanding as infrastructure-level work, not a niche add-on. For you as an end user, the practical payoff would be meeting summaries and call records where you can actually trust which name is next to which quote.

Google's 49th filing in Voice & speech AI we've tracked since May follows work like smarter follow-up questions and fixed-speed audio generation.

Editorial take

The design trade-off here is real and worth naming. By processing audio in short chunks first and then reconciling labels in text space afterward, Google avoids the computational cost of running a single massive model over hours of audio. That is a smart engineering call, but it means the reconciliation network never has access to the actual audio again; it is working from labels alone. If two chunks produce genuinely ambiguous labels because two speakers sound similar or overlap heavily, the reconciler has no acoustic evidence to fall back on.

That's the cost: the accuracy ceiling of stage two is set by the quality of stage one. If the chunked diarization misidentifies a speaker in a segment, the reconciler may faithfully preserve that error across the whole transcript rather than catching it.

Still, for the realistic case where chunked diarization is reasonably accurate but label assignments are just inconsistently numbered, this approach is a proportionate fix. It solves a real, annoying problem without burning compute on a full re-analysis. The filing reads as practical plumbing work, the kind of improvement that makes products noticeably better without earning a press release.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

6 drawing sheets from US 2026/0300624 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.