Google Patents an AI System That Predicts How Heavy Workloads Will Perform Before Running Them
Before a massive computing job even starts, Google's new system tries to tell you how long it will take and what it will cost. The trick: it runs a stripped-down version of the job on minimal hardware, then uses a trained AI model to fill in the rest.
What Google's workload performance estimator actually does
You're an engineer at a company that needs to run a huge data-processing job on Google Cloud. Before you commit thousands of dollars in compute time, you'd love to know: how fast will this actually run? Right now, answering that question can mean spinning up the full cluster of machines just to test, which wastes money and time.
Google's patent describes a system that shortcuts that process. Instead of testing on all the hardware, it runs a cut-down version of your job on a much smaller set of machines. An AI model trained on past jobs then estimates the extra cost of the machines talking to each other at full scale, the part the small test can't capture. Put those two pieces together and you get a predicted performance figure, without ever running the full job.
For you as a user of cloud computing services, this could mean faster, cheaper job planning. Instead of paying to "warm up" a big cluster just to see if your settings are right, the system does most of that math in advance.
… rewriting, by the one or more processors, the workload to run on reduced hardware resources; running, by the one or more processors, the workload on the reduced hardware resources to generate a reduced hardware performance; …
Translation: The system scales down the workload so it can run on smaller hardware to test how fast it goes.
How the ML model fills in what the hardware test leaves out
The patent describes a three-part pipeline for estimating how a large workload will perform on a full cluster of processors or accelerators.
Step 1, Rewrite and shrink. The original workload (a training job for a neural network, for example) is automatically rewritten so it can run on a much smaller number of hardware units. This isn't a full simulation; it's the real job code, just adapted for fewer machines.
Step 2, Run the mini version and collect a profile. That rewritten job runs on the reduced hardware, producing a reduced hardware profile: real timing and resource data from an actual execution, just at small scale. This captures the computation cost accurately.
Step 3, Add the communication cost via ML. At full scale, machines constantly send data back and forth, a cost that the mini run can't reproduce. A machine learning model, trained on previous jobs, takes the original workload as input and predicts what that inter-machine communication overhead will be. The two results are then combined into a single estimated performance profile.
The claim is that this approach costs far less in hardware and memory than running the whole job speculatively, while still producing an accurate enough estimate to be useful for scheduling, budgeting, or configuration decisions.
The workload is also input to a machine learning model trained to determine a communication cost for the workload. The machine learning model outputs the communication cost for the workload.
Translation: An AI calculates how much time machines will waste talking to each other across the network.
What this means for cloud bills and job scheduling
For anyone running workloads on cloud infrastructure, the cost of not knowing how a job will behave is real. Teams often overbuy resources as a buffer, or run expensive trial jobs just to tune settings. A fast, cheap estimation layer could reduce that guesswork and shave meaningful dollars off infrastructure bills.
Google's interest in large-scale workload efficiency fits squarely into the economics of its cloud business, where idle or poorly scheduled compute is lost revenue. This patent sits closer to the infrastructure layer than to consumer products, so you probably won't see a branded feature announcement. But if it works as described, the people who would notice are data engineers and ML teams whose jobs run faster and cost less to plan.
Google files its 44th patent we've tracked in AI training and infrastructure since May, adding to earlier work like one on light in photos and one on picking AI models.
The practical benefit here arrives before you spend anything. Instead of running a massive, expensive computing job only to discover it was misconfigured, you get an accurate cost and performance estimate upfront, using a fraction of the actual resources.
Most people using cloud computing services will never see this working. The benefit shows up as a bill that does not surprise them, or a job that completes instead of failing halfway through. That is a quiet win, but a real one.
The harder question is whether the estimation holds up across very different types of jobs. A model trained on some workloads may not predict well for others, and that gap would only reveal itself when an estimate turns out to be wrong. That is an engineering challenge this patent does not resolve, only sidestep for now.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
5 drawing sheets from US 2026/0300021 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in