IBM Patents AR Captions That Reshape Themselves Around How You Read
IBM has filed a patent for AR captions that don't just sit there on your glasses: they watch how you're reading and change their size, speed, or position on the fly to keep you from falling behind.
What IBM's self-adjusting AR captions actually do
You're at a work presentation and the speaker is moving fast. You're wearing AR glasses that show captions, but the text is scrolling too quickly and you're losing chunks of what's being said.
IBM's patented approach would fix this by watching how you consume captions, not just whether you're looking at them. If you seem to be struggling, the system changes something: the font gets bigger, the words slow down, the caption block moves to a better spot in your field of view. All of this happens automatically, without you needing to pause and dig through settings.
For the roughly 1.5 billion people worldwide who live with some degree of hearing loss, this kind of adaptation could make the difference between following a conversation and guessing at it. Real-time caption adjustment based on actual reading behavior is the core idea here.
… monitoring user processing behavior associated with consumption of the captions; and dynamically modifying, in real time and based on the monitored user processing behavior, at least one presentation parameter of the captions …
Translation: The system watches how you read the captions and instantly adjusts how they look to help you understand better.
How the system tracks reading and rewrites the display
The patent describes a closed-loop system with four main stages that work continuously during a speech event.
- Audio capture and transcription: A microphone picks up a speaker's words and a speech-to-text engine converts them to written captions, as existing tools already do.
- Caption display on AR: Those captions appear on an augmented reality device, like smart glasses, overlaid on the real world in front of the user.
- Behavior monitoring: This is the novel part. The system tracks what it calls "user processing behavior" during caption consumption. That likely means eye-tracking data (where you're looking and for how long), reading pace, and possibly whether you're repeatedly glancing at the same text.
- Dynamic parameter modification: Based on what the system observes, it adjusts at least one presentation parameter in real time. The patent doesn't enumerate every possible parameter, but typical candidates include font size, text scroll speed, caption position, line length, or display contrast.
The feedback loop is continuous: the system keeps monitoring and keeps adjusting throughout the entire presentation, not just at a fixed interval. The goal is matching the display to the user's actual reading capacity in the moment, which can vary with fatigue, speaker pace, or topic complexity.
What this means for deaf and hard-of-hearing users
For people with hearing loss, captions are often the only channel carrying spoken information. When those captions move too fast, appear in a bad spot, or use small text in a cluttered visual field, comprehension breaks down fast. A static caption system designed for the average reader doesn't account for individual differences or changing conditions inside a single conversation.
If this approach works as described, it could make AR-based accessibility tools meaningfully more useful in high-stakes settings like lectures, medical appointments, or legal proceedings. IBM keeps filing on accessibility-driven AI suggests this is part of a broader pattern, though this specific patent is narrowly focused on the display layer rather than the speech recognition underneath it.
IBM's fifth filing we've tracked in the AR glasses race since July builds on earlier applications like one flagging factory defects and one sharing objects by gaze.
Hearing loss is one of the most common disabilities on the planet, and the current state of live captioning is genuinely inconsistent. Captions that are too fast, too small, or displayed in the wrong part of the visual field aren't a minor inconvenience: they erase whole conversations for people who depend on them.
The core idea here, closing the loop between how a person reads and how the display responds, is sound. The hard question is whether behavior monitoring from eye-tracking alone gives the system enough signal to make good real-time decisions, or whether it will misread a moment of distraction as a comprehension problem and start changing things at the wrong time.
The patent is also fairly broad in what it claims, leaving the specific parameters and the specific monitoring signals underspecified. That makes it easier to patent and harder to evaluate. The problem it's attacking is real and large; whether this particular framing delivers a proportionate solution depends entirely on implementation details that aren't in this filing.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
2 drawing sheets from US 2026/0301602 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in