Microsoft Patents Technology That Reads Photos Aloud for Blind Users
Most screen readers can read text on a page, but photos have always been a wall of silence for blind users. Microsoft is patenting a system that identifies objects inside an image and reads them out loud when you move your cursor over them.
How Microsoft's image reader talks about what's in a photo
Imagine you receive an email with a photo attached, or you're browsing a news article full of pictures. If you're blind or have low vision, those images are essentially invisible to your screen reader, which can only read the text around them.
Microsoft's patent describes a tool that uses AI to look at an image and figure out what's in it, both the overall scene (say, "a park on a sunny day") and the specific objects inside it (a dog, a bench, a person). It then connects that information to your cursor or pointer. When you hover over a part of the image, the computer reads out loud what's there.
This goes beyond simple image captions. Instead of one generic description for the whole photo, you'd get different audio details depending on where you point, turning a static image into something you can explore piece by piece.
How the AI classifies objects and triggers audio descriptions
The system works in two main steps: classification and audio playback on demand.
First, a machine learning engine (an AI model trained to recognize visual content) analyzes the image and produces two layers of metadata (structured data attached to the file):
- A scene classification describing the overall image, for example "a city street at night."
- An object classification for specific regions of pixels, identifying individual items like a car, a traffic light, or a person.
Second, the system watches for user input directed at the image. The patent specifically covers three interaction types: hovering a cursor over a region, clicking or selecting a region, or moving the cursor across the image. Any of these actions can trigger audio playback describing whatever object occupies that pixel area.
The result is a kind of interactive audio map of a photo. The AI's object and scene labels become a navigable layer underneath the image, giving blind users a way to explore visual content the same way a sighted person might scan it with their eyes.
What this means for blind and low-vision computer users
Screen readers have existed for decades, but image accessibility has always lagged. Tools like alt text help only when a human has written a description, which most images on the web don't have. An AI-driven, cursor-aware system like this would work on any image automatically, without requiring the original publisher to do anything extra.
For Microsoft, this fits naturally into Windows and Office, where blind users already rely on Narrator and other accessibility tools. A system like this could make photo attachments in Outlook, images in Word documents, or pictures in Edge browsable by touch and audio. Whether or not this specific patent ships as a product, it signals that Microsoft is thinking about accessibility at the image-content level, not just the page-layout level.
This is a genuinely useful accessibility idea, and the cursor-based, object-level approach is meaningfully better than a single static caption. The real question is whether Microsoft ships it broadly across Windows and Office or lets it sit as a narrow feature. Given that the company has been expanding its Narrator and accessibility stack, this one seems like a real candidate to appear in a future OS update.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
17 drawing sheets from US 2026/0227951 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Editorial commentary on a publicly published patent application. Not legal advice.