Disney Patents a System That Cuts Its Own Sports and Event Highlights
Disney has patented a system that listens to a live broadcast, figures out who's being talked about, and automatically cuts video clips around those moments, no editor required. It's the highlight reel that makes itself.
How Disney's auto-clipping system actually works
Every time a broadcaster says a player's name during a live game, a producer somewhere has to decide whether that moment is worth clipping and sending to fans. That decision, repeated hundreds of times an event, is slow, expensive, and inconsistent.
Disney's patent describes a system that handles this automatically. It transcribes the audio from a live event in real time, identifies the names of people, teams, or other key subjects being discussed, and groups the transcript into segments dense with those mentions. Then it cuts the actual video to match those segments, timestamps and all, and sends the clips to users' devices. Clips can be pushed out automatically or held until a fan's personal alert settings trigger them.
The system also accepts feedback to improve its clip selection over time, so it's not locked into a fixed set of rules.
… perform phonetic correction and co-reference resolution of the entities using predetermined phonetic rules and predetermined co-reference rules, respectively …
Translation: It cleans up names and pronouns in the transcript to make sure who or what is being discussed is clear.
Inside Disney's text-segment-to-video-clip pipeline
The system starts by ingesting audio or video from a live event and running it through a transcription engine that attaches a timestamp to every word.
From that timestamped text, the processor identifies entities (named people, teams, places, or other subjects mentioned in the broadcast). It then applies two cleanup passes: phonetic correction (fixing names that were misheard or misspelled by the transcriber, like 'Le Bron' being stitched back to 'LeBron') and co-reference resolution (understanding that 'he scored' refers to the player named two sentences earlier, so the clip window stays coherent).
Next, the text is broken into segments based on preset rules, each segment required to contain at least a minimum number of entity mentions and carrying a clear start and end timestamp. The video is then cut to match those exact timestamps, producing a clip that begins and ends where the relevant conversation does.
- Clips are delivered to user devices for viewing or listening
- Alert notifications can fire automatically or based on user preferences
- A feedback loop lets the system adjust its clipping rules over time
Critically, the patent specifies that all of these steps run contiguously and without human intervention after the media feed is received.
… clipping from the media data the at least one media clip having a begin timestamp and end timestamp corresponding to the begin timestamp and end timestamp of a corresponding one of the text segments …
Translation: It cuts out the video segment based on the exact start and end times of those text blocks.
What this means for live sports and event coverage
For sports and live entertainment, the bottleneck has always been the human editor sitting between the broadcast and the fan's phone. This system removes that bottleneck entirely. If it works as described, highlights could reach fans in seconds rather than minutes, and the volume of clips produced per event could scale far beyond what a production team could manually cut.
Disney's steady investment in live-event automation fits a company that owns ESPN and a massive portfolio of live rights. Fewer editors per event also means lower per-clip production costs, which matters when a network is managing dozens of simultaneous feeds across a sports season.
That makes this Disney's fifth filing in the Enterprise AI patents we cover since June, a group that includes one on auto-tagging its content and one on predicting park equipment failures.
Claim 1 is broad. It covers any automated system that transcribes audio, identifies named entities, groups text into timestamped segments, and cuts video to match, for any event, on any user device. The claim does not restrict itself to sports, to a particular AI model, or even to a specific type of named entity. That scope means it could, if granted, cover a wide swath of automated clipping tools built around transcript-driven segmentation.
The practical consequence is that competitors building real-time highlight tools using a transcribe-then-segment-then-clip pipeline would need to think carefully about this claim. The phonetic correction and co-reference resolution steps add some specificity, but they are described in terms of 'predetermined rules' rather than any particular technology, which keeps the claim's reach wide.
Whether the patent survives examination at that breadth is another question. Transcript-to-clip workflows have existed in broadcast production software for years, and examiners will scrutinize the prior art closely. The feedback loop and alert system add commercial texture but probably not enough novelty on their own to rescue an otherwise thin claim.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
32 drawing sheets from US 2026/0268939 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →