Nvidia · Filed Mar 26, 2025 · Published Oct 1, 2026

Nvidia Patents a Way to Build AI Training Data by Scraping Game Wikis

Training an AI to recognize what it's looking at requires thousands of carefully labeled images, and labeling them by hand is slow and expensive. Nvidia has filed a patent for a system that does much of that work automatically, starting with game wikis.

A screenshot from a video game with multiple square overlays, likely representing detected objects or icons within the game environment. Drawing from patent filing US 2026/0301378 A1.
A screenshot from a video game with multiple square overlays, likely representing detected objects or icons within the game environment.
See all 16 drawings from this filing ↓
Publication number US 2026/0301378 A1
Applicant NVIDIA Corporation
Filing date Mar 26, 2025
Publication date Oct 1, 2026
Inventors Ram Rangan, Deep Shekhar, Ritesh Kumar, Anjul Patney
US classification 382/159
Examiner SUMMERS, GEOFFREY E (Art Unit 2669)
Status when we published Waiting for an examiner (May 1, 2025)
Document 20 claims

How Nvidia turns game wiki images into AI training data

Imagine you're teaching an AI to recognize characters, items, and environments inside a video game. Normally, someone has to sit down and manually tag thousands of screenshots: "this is a sword," "this is the character's health bar," and so on for every single image. That process can take months and cost a lot of money.

Nvidia's patent describes a shortcut. The system browses wiki pages and other web resources built around a specific game, collects images it finds there, and reads the surrounding text for context. Then an AI reads both the images and the text together to figure out what each image shows and attach the right labels automatically.

Once those first images are labeled, the system goes further: it uses them as a reference to find even more similar images and label those too, growing the training library on its own. The goal is a large, labeled dataset built largely without a human doing the tedious tagging work.

From the filing · CLAIM 1
obtaining, based at least on analyzing webpage code corresponding to one or more webpages associated with an interactive application, at least: image data representing one or more first images depicting one or more first objects of the interactive application; and text data representing information associated with the one or more first images; …

Translation: The system pulls pictures and text from web pages about a video game by reading their underlying code.

How the system crawls, labels, and expands image datasets

The system starts by crawling webpages tied to a specific interactive application, most likely a game. It pulls two things from each page: image data (the actual screenshots or artwork) and text data (captions, descriptions, article text sitting near those images).

Both are handed to a multimodal language model (an AI that can read images and text at the same time, similar to how GPT-4o or Gemini work). That model assigns classification labels to the images, for example identifying an object as a specific weapon type, a named character, or a UI element.

Those first labeled images, called seed images in the patent, then power a second step: the system searches for additional images that show objects matching the same classifications. Those new images get added to the dataset too, expanding it well beyond what the web crawl alone found.

The result is a training dataset the system built largely on its own, ready to be used to teach a computer vision model how to identify objects inside that application. The patent frames this as applicable to games specifically but the underlying approach could extend to any domain where labeled images exist on the web.

From the filing · THE ABSTRACT
… crawl network resources (e.g., webpages) that include information associated with an interactive application (e.g., a wiki page for a gaming application) to obtain “seed” images depicting objects from the interactive application, as well as information (e.g., metadata) corresponding to the seed images.

Translation: It automatically browses game wikis to collect starting pictures and data for training AI models.

What this means for AI development costs and speed

Building training data is one of the most expensive and time-consuming parts of developing any computer vision AI. Labeling images manually requires either a large internal team or outsourcing to annotation services, both slow and costly at scale. A system that can bootstrap that process from publicly available web content could cut those costs and timelines considerably.

For Nvidia, which sells the chips and software platforms that power AI training, a growing pile of Nvidia computer vision filings points to a broader interest in lowering the barrier to building AI systems, not just running them. If training data is easier to generate, more developers train more models, and more of that work runs on Nvidia hardware.

Nvidia's 67th filing we've tracked in AI vision since May adds to a run that includes one on cutting redundant frames and one on catching circuit congestion early.

Editorial take

Teaching a machine to recognize objects requires showing it thousands of labeled examples first, and producing those labels means paying people to sit and tag images one by one. For any company building a visual AI system, that work can eat months and significant budget before a single useful model exists.

Nvidia's approach targets that cost directly by pulling from game wikis and similar community-built sites, where players have already done the work of pairing images with descriptions. The system reads those natural pairings to assign labels automatically, then uses them as a foundation for building larger training sets.

The match between problem and approach is credible. The bottleneck in visual AI has always been the sheer volume of human annotation required, and a method that converts existing web content into training data could compress that timeline from months to days across many product categories.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

16 drawing sheets from US 2026/0301378 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.