Sony · Filed Jun 6, 2025 · Published Jul 16, 2026 · verified — real USPTO data

Sony Patent Converts Whispered Speech to Normal Voice Without Text Labels

Whispering in a meeting so you don't disturb anyone nearby is polite, but it wrecks voice recognition software. Sony is working on AI that can fix that.

Sony Patent: Converting Whispers Into Normal Speech With AI — figure from US 2026/0204275 A1
Figure from the official USPTO publication.
Publication number US 2026/0204275 A1
Applicant Sony Group Corporation
Filing date Jun 6, 2025
Publication date Jul 16, 2026
Inventors Junichi REKIMOTO
CPC classification 704/202
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Apr 17, 2026)
Parent application is a National Stage Entry of PCTJP2023035172 (filed 2023-09-27)
Document 21 claims

What Sony's whisper-to-voice AI actually does

Imagine you're on a video call in a crowded coffee shop, whispering so you don't bother the people around you. The problem is that whispered speech sounds completely different from normal speech to a computer: it has no real pitch, no tonal qualities, and most voice software treats it as noise.

Sony's patent describes an AI system trained to bridge that gap. It learns the relationship between whispered speech and normal speech without needing anyone to manually label or transcribe recordings. Feed it a whisper, and it can reconstruct what that speech would sound like if spoken aloud in a full, natural voice.

The system is designed to work across different speakers without being customized for any one person's voice. That's the tricky part Sony says it has solved: teaching the AI to recognize that a whispered 'hello' and a spoken 'hello' are the same thing, even though they sound very different to a microphone.

How the encoder-decoder system bridges whispers and normal speech

The system has two main components working in sequence:

  • Speech-to-unit encoder: This takes a raw audio waveform, whether it's a whisper or normal speech, and compresses it into what Sony calls an acoustic unit. Think of this as a kind of audio fingerprint that captures the linguistic content of what was said, stripped of the vocal qualities (like pitch or breathiness) that differ between whispering and speaking.
  • Unit-to-speech decoder: This takes that acoustic unit and reconstructs it as a full, natural-sounding speech waveform.

The clever part is in how the decoder is trained. Sony uses a technique called self-supervised learning with a Masked Language Model approach (the same family of method that underpins modern AI language tools). Rather than requiring human-labeled transcripts, the model learns by predicting missing parts of audio sequences on its own, using both normal speech and whispered speech simultaneously.

The result is a shared internal representation, the acoustic unit, that treats whispered and spoken versions of the same word as essentially the same thing. The difference between the two vocal styles is "absorbed" in that middle layer, so the decoder can always reconstruct a normal-sounding voice regardless of which style went in.

What this means for meetings, accessibility, and privacy

For remote meetings, this technology could let someone whisper a comment into their microphone and have it come out as a normal voice on the other end. That's useful in open offices, libraries, hospitals, or anywhere that loud speaking is inappropriate. It could also benefit people with voice disorders or conditions that leave them unable to produce full-volume speech.

There's also a privacy angle worth thinking about. A system that reconstructs your voice from a whisper means that even very quiet speech could be captured and amplified in ways that weren't previously possible. That cuts both ways: helpful for accessibility, but worth being aware of as the underlying technology improves and spreads.

Editorial take

This is a genuinely interesting patent from Sony's research division, led by a well-known inventor in human-computer interaction. The self-supervised training approach, which avoids the need for labeled whisper-speech pairs, is the real technical bet here. Whether Sony turns this into a product feature for its audio hardware, conferencing software, or accessibility tools is the open question.

Which company should we read for you?

We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.

Get one Big Tech patent every Sunday

Plain English, intelligent commentary, no hype. Free.

Source. Full patent text and figures from the official USPTO publication PDF.

Editorial commentary on a publicly published patent application. Not legal advice.