Apple Files Patent for Letting AI Servers Work While They Talk to Each Other
When AI servers team up on one job, they waste time waiting on each other. Apple's patent has them send rough, compressed notes ahead so nobody has to pause.
What Apple's overlapping AI data swap does
Two computers split up one big AI job, like two cooks working on the same recipe. At every step, each one stops and waits for the other to hand over its half before either can move on. That waiting is the problem Apple's patent application goes after.
In Apple's filing, each computer sends a rough, compressed version of its work to the other while it is still doing the real calculation. The receiving computer unpacks that rough copy and uses it as a stand-in, so it never has to pause for the exact numbers. The answer comes out slightly approximate, but the machines stay busy.
The payoff is time. The data transfer hides behind the calculation instead of adding to it. The same idea is described for both training an AI model and running one.
… while generating the first portion of the update for the set of features: receiving, by the first processing resource and from a second processing resource, an approximated and compressed representation of a second portion of the update for the set of features …
Translation: One server processes its data while simultaneously receiving compressed data from another server.
How each server estimates its neighbor's work
An AI model is a huge pile of adjustable numbers called weights. When one computer can't hold or crunch them all, the weights are split into portions and spread across several processing resources (servers, or chips inside them). Each one multiplies its portion of the weights by its portion of the features (the numbers flowing through the model, derived from your input) to produce part of an update for the next stage.
Normally each machine then has to wait for the others' parts before moving on. Under claim 1, the first machine instead receives an approximated and compressed version of the second machine's part while it is still doing its own multiplication. It decompresses that version and combines it with its own result to update the features.
The filing and claims add several details:
- In the description, each sender compresses its features and ships them early, and the receiver rebuilds a stand-in for the sender's result.
- Claims 4 and 5 say the compressing and decompressing functions are trained together, and can be Low-Rank Adaptation (LoRA, a method that adds small extra matrices to a model, usually to fine-tune language models).
- The same pattern is claimed for training a model (claim 20). The description extends it to any number of machines and to transformer models (the design used by many large language models).
By allowing approximation, compression, communication, and computation to occur in parallel, the system may reduce communication overhead that typically causes idle compute periods.
Translation: Doing multiple tasks at the same time stops servers from wasting time waiting for each other.
Why idle AI servers cost real money
You'll never see this one on a settings screen, but you could feel it as speed. If AI servers spend less time waiting on each other, answers can arrive sooner and the same hardware can do more work. The filing says this can raise sustained utilization, improve responsiveness, and lower overall system cost, since the links between machines don't have to be as fast.
The patent describes clusters of servers that train and run models for devices like your phone, so any benefit would land mostly behind the scenes. The trade is accuracy: the filing says results can differ from the full-precision answer in exchange for less communication. It is only a patent application, so nothing here promises a product.
Apple files its eighth patent we've tracked in AI chip competition since August, adding to earlier work like one on cutting chip memory use and one on keeping idle chips current.
Claim 1 reaches further than the headline idea suggests. It never says how the data gets shrunk. It only requires that one machine receive an approximated, compressed piece of another machine's work while it is still calculating its own piece, unpack it, and fold it in.
The specific tricks, a matched pair of functions trained together and the LoRA technique, appear only in later claims (4 and 5). So the opening claim, if granted, would reach any compression method that fits that pattern. The same wording repeats for running a model, training one, a computer system, and stored software.
The catch is the word "approximated." The filing admits the results differ slightly from the full-precision answer, and the claims say nothing about how much error is acceptable. Whether that trade pays off is a question for Apple's engineers, and the claims leave them plenty of room.
Get our take in your Top Stories
Liked this breakdown? Add Patentlyze as a preferred source on Google, and our plain-English take shows up more often in your Top Stories the next time Apple patent news breaks.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
17 drawing sheets from US 2026/0311057 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in