Microsoft · Filed May 15, 2026 · Published Sep 17, 2026 · verified — real USPTO data

Microsoft Patents a Way to Send Your Voice to AI Without Exposing Who You Are

Every time you talk to a voice assistant or transcription service, your words travel across the internet with something extra attached: the unique fingerprint of your voice. Microsoft is now patenting a way to strip that fingerprint out before your speech leaves your device.

A secure voice system encodes original speech, transmits it over a network, and decodes it into a secure speech signal. Drawing from patent filing US 2026/0279371 A1.
A secure voice system encodes original speech, transmits it over a network, and decodes it into a secure speech signal.
See all 6 drawings from this filing ↓
Publication number US 2026/0279371 A1
Applicant Microsoft Technology Licensing, LLC
Filing date May 15, 2026
Publication date Sep 17, 2026
Inventors Dushyant SHARMA, Patrick A. NAYLOR, Chandramouli Shama SASTRY
CPC classification 704/500
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Jun 11, 2026)
Parent application is a Continuation of 18532900 (filed 2023-12-07)
Document 20 claims

How Microsoft hides your identity inside a voice call

When you use a voice assistant, your phone or computer doesn't just send what you said. It sends a full recording of your voice, complete with the personal characteristics that make it identifiably yours. Those characteristics can be used to figure out who you are, even if nobody writes down your words.

Microsoft's new patent describes a system that separates your voice into two parts before sending it anywhere: the content (what you said) and the speaker identity (the unique sound of you). The system keeps the identity on your device and sends only the content across the network to a speech recognition system. A digital watermark is also baked in to mark the transmission as legitimate.

The goal is to let AI transcription and voice assistants do their job without ever getting access to the personal information that could identify you. Your words travel; your vocal fingerprint stays home.

From the filing · CLAIM 1
… encoding, by the trained neural network, the speech signal into a content embedding to obtain an encoded signal, the content embedding including less of the speaker information than the speech signal …

Translation: The system scrubs identifying details out of your voice data before sending it.

How the encoder separates content from speaker identity

The system works in two machine-learning stages happening inside an encoder (software that packages your voice for transmission).

First, the encoder analyzes the raw speech signal and extracts a speaker embedding, a compact mathematical snapshot of the personal characteristics of your voice (things like pitch, resonance, and speaking rhythm that remain consistent across different words you say). Think of it as a short numerical description of what makes your voice yours.

Second, that speaker embedding is fed back into the process as a reference point, and the encoder uses it to generate a content embedding, a different numerical package that carries just the words, with the speaker characteristics actively minimized. The training process essentially teaches the network to subtract identity information.

Before the content embedding is sent across the network, the system applies a neural watermark, an invisible digital tag woven into the data using a neural network. This watermark can later be checked to confirm the signal is authentic and hasn't been tampered with.

The encoded signal is then sent to a remote automatic speech recognition (ASR) system, the AI engine that turns speech into text, which receives content it can transcribe without having access to who said it.

From the filing · THE ABSTRACT
The content component of the voice signal is processed, using machine learning and based at least on the speaker embedding, to generate a content embedding having minimized speaker information.

Translation: Machine learning strips away your unique vocal traits while keeping your actual words.

What this means for voice AI and personal privacy

Voice data is unusually sensitive. Unlike a password, you can't change your voice if it leaks. Services that process speech for transcription, translation, or commands routinely receive full audio recordings, and those recordings can be stored, analyzed, or exposed. The problem is real and not hypothetical: researchers have demonstrated speaker identification from short audio clips with high accuracy.

For everyday users, this kind of system could mean that dictating a message or asking your assistant a question no longer hands a third-party server a sample of your biometric data. That matters most in regulated industries, like healthcare or legal work, where voice transcription tools are common but voice data is especially sensitive. The neural watermark layer also hints at a broader interest in authenticated, tamper-evident AI communication.

This is the sixth Microsoft filing we've tracked since August in our on-device AI privacy watch, following applications like personalized ads without your identity and hiding mouse clicks from your PC.

Editorial take

Speech carries more personal information than most people realize. Your voice reveals your identity as reliably as a fingerprint, which means every recorded meeting, every transcribed call, every voice command sent to a cloud server is also a transfer of something that cannot be changed if it leaks.

Microsoft's patent attacks that problem at the right moment: before the voice leaves your device, not after. The system strips out the identifying layer of a voice signal and sends only the spoken content onward, keeping what you said separate from who said it.

Whether this reaches everyday users or stays tucked inside corporate compliance tools is an open question. Meeting transcription software is the most likely home for it, and that audience, executives, legal teams, healthcare workers, may actually be where the demand for this kind of protection is highest and the consequences of getting it wrong are most costly.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

6 drawing sheets from US 2026/0279371 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.