IBM Patents a System That Builds Training Computers on Borrowed Data and Tells Developers What to Fix
Building a predictive AI model when your company's own data is thin or messy is a common and expensive problem. IBM's new patent tries to solve that by swapping in synthetic data and then nudging developers in the right direction based on how the models actually perform.
How IBM wants to close the AI training data gap
Most companies that want to build AI models run into a basic problem: their own collected data is too sparse, too messy, or too sensitive to train on directly. That gap often kills AI projects before they start.
IBM's patent describes a system that looks at what data a company does have, picks relevant synthetic datasets (data that is artificially generated rather than collected from real events) to fill the gap, and uses that synthetic data to train a set of predictive models. It then tests those models against a held-out portion of the same synthetic data to see how well they perform.
The interesting part is what happens next. Instead of just spitting out accuracy numbers, the system shows the developer a prompt on their screen: a specific suggestion about which model to keep, retrain, or discard. You get a guided decision, not a spreadsheet to interpret on your own.
… presenting prompting data on a displayed user interface of a developer user in dependence on result data resulting from the testing, the prompting data prompting the developer user to direct action with respect to one or more model of the set of predictive models.
Translation: The system shows the programmer specific instructions on their screen about how to fix or improve their AI models.
How the system trains, tests, and nudges developers
The patent describes a pipeline with five linked steps:
- Examining the enterprise dataset: The system looks at what data the company already has, assessing its scope and gaps.
- Selecting synthetic datasets: Based on that examination, it picks one or more synthetic datasets (pre-built or generated data that mirrors the real-world domain but was not collected by the company itself).
- Training predictive models: A set of models is trained using the synthetic data, not the company's own records.
- Testing with holdout data: A reserved slice of the synthetic data (the "holdout" set, meaning data the models did not train on) is used to measure performance.
- Presenting prompting data: Results are shown to a developer through a UI that includes explicit prompts about what action to take on each model.
The claim is broad in that it does not specify what kind of predictive model is being trained, what the synthetic data looks like, or how the prompts are generated from results. The core idea is the loop: real enterprise data informs synthetic data selection, synthetic data trains the models, and the system closes the loop by directing human attention to what needs fixing.
… selecting one or more synthetic dataset in dependence on the examining, the one or more synthetic dataset including data other than data collected by the enterprise …
Translation: The software looks at a company's private information and then picks outside, artificial data to help train its AI.
What this means for companies building AI with limited data
For companies that want to use AI but do not have enough clean internal data to train on, this kind of pipeline could lower the barrier meaningfully. The developer no longer has to know which model architecture to pick or how to read loss curves; the system surfaces an action item. That guided step is especially valuable in enterprise settings where the person running the ML tool may not be a machine learning specialist.
The filing sits in a long stream of IBM research around enterprise AI tooling, and it reflects a broader industry push to make model development accessible to developers who are not data scientists. Readers interested in how companies like IBM are shaping that direction can follow coverage of Big Tech patent news as synthetic-data techniques move from research labs into production pipelines.
Claim 1 is written at a high level of abstraction: it covers any method that examines enterprise data, selects synthetic data, trains models, tests them, and shows a developer a prompt. There is no requirement for a specific algorithm, data type, or prompting strategy. That breadth means the claim could, if granted, reach a wide range of tools that use synthetic data to assist model selection and feed results into a developer-facing interface. The practical question is whether the prior art in automated ML (AutoML) platforms and synthetic data generation tools already describes this loop, which would narrow or cancel the claim during examination. As written, this is an ambitious claim, and its fate will depend heavily on how the examiner defines the state of the art in interactive AutoML.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
13 drawing sheets from US 2026/0236847 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →