Google Patents a Way to Label Who Said What in Live Conversation Transcripts
When two people are talking and a transcript appears on screen, the hardest part isn't writing down the words, it's knowing whose words they are. Google's new patent tackles that by listening to where voices come from, not who they belong to.
How Google's directional audio transcript labeling works
You're sitting across from someone in a meeting and a live transcript is running on your phone or laptop. The words appear fast enough, but every line just shows up in one undifferentiated block. Who said what? You have to read back through the whole thing to figure it out.
Google's patent describes a system that solves this with two microphones. By measuring the tiny difference in when a sound reaches the first microphone versus the second, the system can figure out which direction each voice is coming from. Voice from the left gets labeled as Person A; voice from the right gets labeled as Person B.
The result is a transcript that looks more like a screenplay, with each person's lines clearly attributed. No login, no voice enrollment, no camera required. The system works from geometry alone, which means it could run on anything with two microphones.
… estimating, by the computing device and based on the first and second audio signals, a time delay in respective arrival times for the speech input at the first and second audio input devices …
Translation: The system measures how much earlier sound hits one microphone than the other.
How two microphones pinpoint each speaker's position
The system takes in audio from two separate microphones placed at different positions. As each person speaks, their voice reaches the two microphones at slightly different times, because sound has to travel a tiny bit farther to reach one mic than the other.
The patent calls this a time delay of arrival (TDOA), the microscopic gap (often just fractions of a millisecond) between when a sound hits microphone one versus microphone two. From that gap, the system estimates the direction of the sound source, essentially triangulating where in space the speaker is sitting relative to the two mics.
Once it knows that Person A's voice consistently arrives from, say, the left, and Person B's arrives from the right, it associates each chunk of the speech-to-text transcript with the correct directional source. The final display is a labeled transcript:
- Each line is tagged with the speaker it belongs to
- Labels update in real time as the conversation progresses
- No pre-enrollment of voices is needed; direction alone does the work
The patent is explicit that this uses two audio input devices, not one. That two-microphone minimum is what makes the direction-finding math possible.
… displaying based on the associating of the respective portions, a modified speech-to-text transcript of the conversation that labels the respective portions of the speech-to-text transcript associated with the respective participants …
Translation: The screen shows a text transcript with the speakers' names clearly marked next to their words.
What this means for live captioning and transcription apps
Live transcription tools already exist in Google Meet, on Android, and in third-party apps. But speaker diarization (the technical term for labeling who said what) usually depends on voice-recognition models trained on your voice, or on a logged-in meeting participant list. That requires setup, accounts, and data. This approach skips all of that and works from the physical positions of the speakers.
For accessibility tools specifically, clear speaker labels make transcripts far more useful for people who are deaf or hard of hearing. A wall of text is much harder to follow than a back-and-forth dialogue view. Google's interest in on-device accessibility and real-time captioning suggests this could eventually show up in Android's Live Transcribe or a future version of Google's wearable audio products.
Google's 795th filing we've tracked since May in our Google coverage continues a pattern seen in the energy outage predictor and the smart home action suggester.
Claim 1 is broad in a way that should raise eyebrows. It doesn't specify what kind of device the two microphones sit on, how far apart they need to be, or what speech-to-text engine does the transcribing. Any system that receives two audio signals, estimates direction from time delay, and outputs a labeled transcript could fall inside this claim as written.
That breadth cuts both ways. A broad claim is harder to get through the patent office, but if granted, it would cover a wide range of implementations: a phone with two mics, a laptop, a dedicated transcription device, possibly even a pair of earbuds in each ear of two different people. That's a lot of territory to hold.
The practical value of the underlying idea is real. Speaker attribution in live transcripts is a genuinely unsolved problem for most casual users. Whether this specific claim survives examination largely intact is another question, but the direction (so to speak) is clearly toward making transcription tools work without asking you to do any setup first.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
12 drawing sheets from US 2026/0301745 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in