Google Patents a System That Displays AI-Generated Info Cards During Video Playback
Google is working on a system that reads a video as it plays and automatically surfaces a pop-up information card about whatever person, place, or thing just appeared on screen, before you even think to pause and search.
What Google's mid-video entity cards actually do
Right now, when something in a video sparks your curiosity, a celebrity you half-recognize, a product someone's using, a landmark in the background, your only option is to stop the video, switch apps, and search for it yourself. That friction causes a lot of people to just forget about it.
Google's patent describes a system that watches the video alongside you and predicts what you're likely to want to look up. When the right moment comes, a small card appears on screen with descriptive information about that entity, a person, brand, location, or object, while the video keeps playing.
The key detail is that the AI doesn't just recognize what's in the frame: it's trained to predict what a viewer would actually search for, which is a more selective and useful filter than simply labeling everything it sees.
… in response to a first entity occurring in the video, providing, for presentation on the display device while the video is being played, a first entity card which includes descriptive content relating to the first entity …
Translation: The system automatically shows an information box about a person or object while you are watching a video.
How the model predicts what you'd search for next
The system uses one or more machine-learned models (AI trained on large datasets) to process a video as it plays. The models analyze the content frame by frame, identifying what Google calls entities, people, places, brands, products, or concepts that appear or are referenced in the video.
Critically, the model's job isn't just object recognition. It's predicting which entities a viewer is likely to search for, a more refined goal that filters out background noise and focuses on things that carry genuine curiosity value. That predicted entity then triggers the generation of an entity card, a structured information panel displayed on the same screen while playback continues.
The patent describes the card as containing descriptive content relating to the entity, though it doesn't specify exactly what form that takes. Likely candidates include names, short bios, links, or Knowledge Graph-style summaries.
- Video is processed by AI models in real time or near-real time during playback
- Models output a ranked prediction of entities worth surfacing
- A card is generated and displayed on-screen without interrupting the video
- The card is tied to the moment the entity appears in the video
The first entity is identified by one or more machine-learned models based on the one or more machine-learned models processing, as an input, the video, and predicting, as an output, the first entity as an entity likely to be searched for.
Translation: AI analyzes the video to guess which people or items you might want to look up later.
What this means for YouTube's search-and-watch loop
For anyone who watches video on YouTube or Google TV, this would change the basic rhythm of how you consume content. Instead of the watch-pause-search-return loop that breaks viewing, relevant information would arrive at the right moment, on its own. That's a meaningful quality-of-life change, and it also keeps you inside Google's ecosystem rather than jumping to another app.
From a business angle, entity cards are an obvious surface for rich results, affiliate-style product links, or ad-adjacent placements, though the patent doesn't say any of that explicitly. The filing sits squarely in the broader race to make video as information-dense as text, an area where Big Tech patent news has tracked a steady stream of AI-driven video-understanding filings from Google and its rivals over the past two years.
The underlying AI capability described here, predicting search intent from video rather than merely recognizing objects, is the part that takes real engineering work, and Google is probably closer to shipping it than most companies would be. The patent is pure software running on existing video infrastructure, so there's no new hardware gate to clear; the shortest route to a product is a server-side model update and a YouTube UI change. The gap between this filing and a live feature could realistically be measured in months rather than years, assuming the prediction quality is high enough that cards feel helpful rather than intrusive.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
13 drawing sheets from US 2026/0236531 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →