Microsoft · Filed Jun 8, 2026 · Published Oct 1, 2026 · verified — real USPTO data

Microsoft Patents an AI That Pulls One Voice Out of a Noisy Video Call

Every video call has the same problem: echoes, crosstalk, and ambient chatter fighting for the microphone. Microsoft is patenting an AI model that handles all of it in one pass, keeping only the voice of whoever you actually want to hear.

Different scenarios of personalized speech extraction from noisy audio, isolating specific voices from background noise and other speakers. Drawing from patent filing US 2026/0301758 A1.
Different scenarios of personalized speech extraction from noisy audio, isolating specific voices from background noise and other speakers.
See all 11 drawings from this filing ↓
Publication number US 2026/0301758 A1
Applicant Microsoft Technology Licensing, LLC
Filing date Jun 8, 2026
Publication date Oct 1, 2026
Inventors Sefik Emre ESKIMEZ, Takuya YOSHIOKA, Huaming WANG, Alex Chenzhi JU, Min TANG, Tanel PÄRNAMAA
CPC classification 704/202
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Jun 27, 2026)
Parent application is a Continuation of 18172017 (filed 2023-02-21)
Document 20 claims

How Microsoft's call-cleanup AI decides whose voice to keep

Ever been on a video call where you could hear the other person's audio bouncing back at you, plus someone talking in the background on your end? Both problems are usually handled by separate software systems that don't always work well together.

Microsoft's patent describes a single AI model that tackles both at the same time. You tell the system whose voice you want, and it strips out echoes (your own speakers bleeding back into the microphone) and other people's voices in one step, handing you a clean audio track of just your target speaker.

The key twist is the personalization piece. The model learns what your target speaker sounds like and uses that as a guide, so it isn't just reducing noise generically. It's actively looking for one specific voice and discarding everything else.

From the filing · CLAIM 1
a self-attention alignment block configured to use attention to determine a soft-alignment of first features of the first embeddings and second features of the second embeddings, the self-attention alignment block outputting attention weights associated with the soft-alignment; …

Translation: The AI compares different audio streams to figure out which sounds belong together.

Inside the two-stage LSTM pipeline that handles echo then noise

The system takes three inputs: the audio coming from the far end of a call (what the remote participant's speaker is playing), the audio captured by the near-end microphone (which contains the target speaker, background talkers, and the echo of that far-end audio), and a voice profile representing the target speaker's acoustic characteristics.

Those inputs are each encoded into embeddings (numerical summaries that capture audio features) and fed into a machine learning model with two sequential stages.

  • A self-attention alignment block (a component that finds where the far-end and near-end audio overlap in time) computes "attention weights," essentially a map of which moments in the far-end signal match the echo in the near-end recording.
  • A first group of LSTM blocks (Long Short-Term Memory networks, a type of AI that handles sequences by remembering earlier context) uses that map to remove the echo signal.
  • A second group of LSTM blocks then takes the echo-cleaned audio plus the target speaker's voice profile and suppresses any remaining interfering speakers, outputting only the target voice.

Running echo cancellation first and noise suppression second in a single trained model lets each stage inform the other, rather than running as separate, uncoordinated filters.

From the filing · THE ABSTRACT
The machine learning model trained to analyze the far-end signal and the near-end signal to perform personalized noise suppression (PNS) to remove speech from one or more interfering speakers and acoustic echo cancellation (AEC) to remove echoes.

Translation: The system uses machine learning to strip out background noise and echoes from a call.

What this means for Teams calls in crowded offices

For anyone on a video call in an open office or shared space, this is a direct quality-of-life problem. Today's call software often handles echo and background noise through separate components that can interfere with each other or require the user to do nothing special except hope for the best. A unified model that is also personalized to a specific speaker's voice could mean cleaner audio with fewer artifacts.

For Microsoft, which runs Teams across hundreds of millions of users, even a modest improvement in call audio quality at scale is a meaningful product advantage. a growing pile of Microsoft audio and speech-AI filings suggests this is an area the company is treating as infrastructure, not a niche add-on.

Microsoft's 14th filing we've tracked in Voice & speech AI since May follows earlier work on prepping calls for recognition and blending phone and internet for text-to-speech.

Editorial take

The building blocks for this are already in every laptop and phone: a microphone, a speaker, and a voice call. No new hardware is required. The software just needs a brief sample of whose voice to keep, then filters out everyone else along with any echo.

The missing piece the document doesn't address is speed. Audio that arrives even a fraction of a second late makes a call feel broken, and running this kind of filtering in real time is demanding work for an ordinary device.

The level of architectural detail here suggests something that has been built and tested rather than sketched on a whiteboard, which puts this closer to a product than most patents. Whether it ships depends almost entirely on whether it can run in the background without killing your battery.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

11 drawing sheets from US 2026/0301758 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.