Google Patents a Way to Search the Web Using What's Inside a Video
You're watching a video and want to search for that jacket someone's wearing, but your only tool is a text box. Google has patented a system that reads the contents of multimedia and silently rewrites your search query to match what's actually on screen.
How Google wants to turn video objects into search queries
Right now, searching for something you spot inside a video means pausing, guessing what words to type, and hoping Google understands what you mean. The search engine has no idea what's in the video you're watching, so you're on your own.
Google's new patent describes a system that changes this. It scans a photo or video for recognizable objects and details, then uses those details to improve your search query automatically. So if you type "buy this" while watching a clip, the system might rewrite that behind the scenes as "blue suede Chelsea boots women's size 8" based on what it detected in the frame.
The idea is that your search becomes about what you're actually looking at, not just the words you could think to type. Google ranks several possible rewrites, picks the best one, and sends that to its search engine instead of your original, vague question.
… extracting one or more entities that appear in the multimedia content item within a predetermined range of time relative to the current playback time; generating a rewritten query by combining the one or more terms of the query with the one or more entities …
Translation: It grabs what is currently visible on screen and merges it with your text search.
How the system extracts entities and rewrites your query
The patent describes a pipeline with several connected steps:
- Entity extraction: The system analyzes multimedia content and pulls out "entities," meaning structured data points that describe what's in the video or image. These might include object types, colors, brands, text visible on screen, or other attributes.
- Query rewrite generation: It takes your original search query and combines it with those extracted entities to generate multiple candidate rewrites. Each candidate is a more specific, context-rich version of what you actually typed.
- Scoring and ranking: Each candidate rewrite gets a score, likely based on how relevant and likely to return good results it is. The candidates are ranked by those scores.
- Search execution: The top-ranked rewrite replaces your original query and gets sent to Google's search engine. The results you see come from that improved query, not the one you typed.
The patent frames this as a general technique that works across multimedia types. The "entities" are essentially machine-readable labels that describe the content, and the rewriting logic bridges the gap between what you said and what you meant.
… providing for display, responsive to the query related to the multimedia content, a result set from the search engine based on the rewritten query.
Translation: It shows search results on your screen while the video keeps playing normally.
What this means for how you find products from videos
For everyday users, this would remove one of the most common frustrations with video browsing: you see something you want, but turning it into a search that actually works takes real effort. A system like this could let you click "search this" mid-video and get back results for the exact item on screen without any extra typing.
Google keeps filing on multimodal search and query understanding, and this patent fits that pattern. From a product angle, the most obvious home for this technology is YouTube or Google Lens, where users already bounce between watching content and wanting to buy or learn about what they see. If this works as described, it could also shift how advertisers think about contextual targeting inside video.
Google's 817th filing in our Google coverage since May builds on earlier applications like the two-model ranking system and the AI search training method.
The core idea here is straightforward software: take what the computer sees, combine it with what the user typed, and produce a better search query. No new hardware required, no specialized chip, no camera module that doesn't already exist on every phone Google already supports.
That makes this one of the shorter routes from patent to product in Google's filing catalog. The hard parts, object detection in video and query rewriting, are things Google already does at scale in separate products. The patent is really about combining them into one pipeline and putting the result in front of a user's search.
The real question is accuracy. A rewrite that misidentifies a jacket as a shirt, or a lamp brand as a generic noun, produces results worse than what you'd have typed yourself. Getting the ranking of candidate rewrites right is where this either becomes useful or disappears.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
5 drawing sheets from US 2026/0300303 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in