IBM · Filed Feb 17, 2025 · Published Aug 20, 2026 · verified — real USPTO data

IBM Patents a System That Delays Video Calls to Sync Translated Captions

Translated captions on video calls almost always lag behind the speaker, showing up after the moment has passed. IBM's new patent flips that around by holding the video itself until the captions are ready.

Videotelephony terminals connected to a central controller equipped with translation and captioning modules. Drawing from patent filing US 2026/0246893 A1.
Videotelephony terminals connected to a central controller equipped with translation and captioning modules.
See all 3 drawings from this filing ↓
Publication number US 2026/0246893 A1
Applicant International Business Machines Corporation
Filing date Feb 17, 2025
Publication date Aug 20, 2026
Inventors Guang Han Sui, Peng Hui Jiang, Jun Su, ZHI LI GUAN
CPC classification 348/14.08
Grant likelihood Medium
Examiner JONES, CARISSA ANNE (Art Unit 2691)
Status Non Final Action Mailed (Aug 10, 2026)
Document 20 claims

How IBM's caption-delay fix works on video calls

Ever been on a video call where the captions appeared several seconds after the person stopped talking? That gap makes captions nearly useless for anyone relying on them to follow a conversation in a different language.

IBM's patent describes a system that solves this by deliberately holding back the video feed while translation and captioning catch up. Instead of rushing captions onto a live stream and hoping they arrive on time, the system treats each short segment of the call like a small package: translate it, caption it, then release it to the viewer. The result is that captions and video arrive together, in sync.

This matters most for international business calls or any situation where two people are speaking different languages. Instead of watching someone's mouth move and reading a caption that belongs to something said three sentences ago, you get a picture that matches the words underneath it.

From the filing · CLAIM 1
… delaying the videotelephony feed until the translating and captioning are completed for each segment of the videotelephony feed; and displaying the delayed videotelephony feed with captions to the second videotelephony terminal.

Translation: The system intentionally pauses the video stream so that the translated subtitles can catch up before the viewer sees them.

How the translation and captioning pipeline times the delay

The patent describes a videotelephony controller sitting between the two callers, essentially acting as a real-time relay station with two jobs.

First, a translation module converts speech from the caller's language into the viewer's language. Second, a captioning module burns those translated words onto the video frame. The key step is timing: the system delays the video feed for each segment until both translation and captioning are finished before passing the segment on.

  • Receiving: The system ingests the live video feed from the first caller's terminal.
  • Translating: Each segment of speech is converted to the target language.
  • Captioning: The translation is overlaid onto the video frames for that segment.
  • Releasing: Only after both steps complete does the system send the segment to the second caller's screen.

The patent is written around a system claim covering memory, a processor, and the software pipeline that orchestrates these steps. It doesn't specify any particular translation engine or captioning format, meaning the architecture is designed to be modular.

From the filing · THE ABSTRACT
… routing the videotelephony feed through a videotelephony controller comprising a translation module and a captioning module, translating, by the translating module, communication from a language used in the videotelephony feed into a language of a second videotelephony terminal …

Translation: The software routes your video call through a special controller that automatically translates the audio into another language.

What this means for cross-language video calls

For anyone who has sat through a multilingual video meeting trying to piece together captions that arrive out of order, this is a straightforward quality-of-life fix. The delay is the point: a fraction-of-a-second hold on the outgoing video buys enough time to produce captions that actually correspond to what's on screen. That's genuinely useful for accessibility and for international collaboration, two areas where video-call platforms compete hard.

The approach is narrow enough that it sits comfortably alongside the stream of new Big Tech patents targeting video communication infrastructure, a category that has seen steady filing activity since remote work normalized cross-timezone and cross-language calls.

Editorial take

Claim 1 covers the whole pipeline at a fairly high level: receive a feed, translate it, caption it, delay it, display it. That breadth means the claim, if granted, could touch any video-call system that adds translation-based captions through a deliberate delay mechanism, which is a wide net. The practical effect is that a competitor building an identical sync-by-delay approach would need to design around this. What the claim does not do is nail down any particular translation method or delay duration, so workarounds exist at the implementation level. The filing is more about owning the architectural idea than a specific technical trick.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

3 drawing sheets from US 2026/0246893 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.