Sony Patent Converts Whispered Speech to Normal Voice Without Text Labels
Whispering in a meeting so you don't disturb anyone nearby is polite, but it wrecks voice recognition software. Sony is working on AI that can fix that.
What Sony's whisper-to-voice AI actually does
Imagine you're on a video call in a crowded coffee shop, whispering so you don't bother the people around you. The problem is that whispered speech sounds completely different from normal speech to a computer: it has no real pitch, no tonal qualities, and most voice software treats it as noise.
Sony's patent describes an AI system trained to bridge that gap. It learns the relationship between whispered speech and normal speech without needing anyone to manually label or transcribe recordings. Feed it a whisper, and it can reconstruct what that speech would sound like if spoken aloud in a full, natural voice.
The system is designed to work across different speakers without being customized for any one person's voice. That's the tricky part Sony says it has solved: teaching the AI to recognize that a whispered 'hello' and a spoken 'hello' are the same thing, even though they sound very different to a microphone.
How the encoder-decoder system bridges whispers and normal speech
The system has two main components working in sequence:
- Speech-to-unit encoder: This takes a raw audio waveform, whether it's a whisper or normal speech, and compresses it into what Sony calls an acoustic unit. Think of this as a kind of audio fingerprint that captures the linguistic content of what was said, stripped of the vocal qualities (like pitch or breathiness) that differ between whispering and speaking.
- Unit-to-speech decoder: This takes that acoustic unit and reconstructs it as a full, natural-sounding speech waveform.
The clever part is in how the decoder is trained. Sony uses a technique called self-supervised learning with a Masked Language Model approach (the same family of method that underpins modern AI language tools). Rather than requiring human-labeled transcripts, the model learns by predicting missing parts of audio sequences on its own, using both normal speech and whispered speech simultaneously.
The result is a shared internal representation, the acoustic unit, that treats whispered and spoken versions of the same word as essentially the same thing. The difference between the two vocal styles is "absorbed" in that middle layer, so the decoder can always reconstruct a normal-sounding voice regardless of which style went in.
What this means for meetings, accessibility, and privacy
For remote meetings, this technology could let someone whisper a comment into their microphone and have it come out as a normal voice on the other end. That's useful in open offices, libraries, hospitals, or anywhere that loud speaking is inappropriate. It could also benefit people with voice disorders or conditions that leave them unable to produce full-volume speech.
There's also a privacy angle worth thinking about. A system that reconstructs your voice from a whisper means that even very quiet speech could be captured and amplified in ways that weren't previously possible. That cuts both ways: helpful for accessibility, but worth being aware of as the underlying technology improves and spreads.
This is a genuinely interesting patent from Sony's research division, led by a well-known inventor in human-computer interaction. The self-supervised training approach, which avoids the need for labeled whisper-speech pairs, is the real technical bet here. Whether Sony turns this into a product feature for its audio hardware, conferencing software, or accessibility tools is the open question.
Which company should we read for you?
We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.
Get one Big Tech patent every Sunday
Plain English, intelligent commentary, no hype. Free.
Editorial commentary on a publicly published patent application. Not legal advice.