Apple · Filed Mar 25, 2026 · Published Oct 8, 2026

Apple Files Patent for Letting AI Servers Work While They Talk to Each Other

When AI servers team up on one job, they waste time waiting on each other. Apple's patent has them send rough, compressed notes ahead so nobody has to pause.

A smartphone connects to paired servers designed to share workloads over a network without pausing calculations. Drawing from patent filing US 2026/0311057 A1.
A smartphone connects to paired servers designed to share workloads over a network without pausing calculations.
See all 17 drawings from this filing ↓
Publication number US 2026/0311057 A1
Applicant Apple Inc.
Filing date Mar 25, 2026
Publication date Oct 8, 2026
Inventors Fernando A. MUJICA
US classification 706/12
Status when we published Waiting for an examiner (Apr 30, 2026)
Parent application Claims priority from a provisional application 63782571 (filed 2025-04-02)
Document 20 claims

What Apple's overlapping AI data swap does

Two computers split up one big AI job, like two cooks working on the same recipe. At every step, each one stops and waits for the other to hand over its half before either can move on. That waiting is the problem Apple's patent application goes after.

In Apple's filing, each computer sends a rough, compressed version of its work to the other while it is still doing the real calculation. The receiving computer unpacks that rough copy and uses it as a stand-in, so it never has to pause for the exact numbers. The answer comes out slightly approximate, but the machines stay busy.

The payoff is time. The data transfer hides behind the calculation instead of adding to it. The same idea is described for both training an AI model and running one.

From the filing · CLAIM 1
… while generating the first portion of the update for the set of features: receiving, by the first processing resource and from a second processing resource, an approximated and compressed representation of a second portion of the update for the set of features …

Translation: One server processes its data while simultaneously receiving compressed data from another server.

How each server estimates its neighbor's work

An AI model is a huge pile of adjustable numbers called weights. When one computer can't hold or crunch them all, the weights are split into portions and spread across several processing resources (servers, or chips inside them). Each one multiplies its portion of the weights by its portion of the features (the numbers flowing through the model, derived from your input) to produce part of an update for the next stage.

Normally each machine then has to wait for the others' parts before moving on. Under claim 1, the first machine instead receives an approximated and compressed version of the second machine's part while it is still doing its own multiplication. It decompresses that version and combines it with its own result to update the features.

The filing and claims add several details:

  • In the description, each sender compresses its features and ships them early, and the receiver rebuilds a stand-in for the sender's result.
  • Claims 4 and 5 say the compressing and decompressing functions are trained together, and can be Low-Rank Adaptation (LoRA, a method that adds small extra matrices to a model, usually to fine-tune language models).
  • The same pattern is claimed for training a model (claim 20). The description extends it to any number of machines and to transformer models (the design used by many large language models).
From the filing · THE ABSTRACT
By allowing approximation, compression, communication, and computation to occur in parallel, the system may reduce communication overhead that typically causes idle compute periods.

Translation: Doing multiple tasks at the same time stops servers from wasting time waiting for each other.

Why idle AI servers cost real money

You'll never see this one on a settings screen, but you could feel it as speed. If AI servers spend less time waiting on each other, answers can arrive sooner and the same hardware can do more work. The filing says this can raise sustained utilization, improve responsiveness, and lower overall system cost, since the links between machines don't have to be as fast.

The patent describes clusters of servers that train and run models for devices like your phone, so any benefit would land mostly behind the scenes. The trade is accuracy: the filing says results can differ from the full-precision answer in exchange for less communication. It is only a patent application, so nothing here promises a product.

Apple files its eighth patent we've tracked in AI chip competition since August, adding to earlier work like one on cutting chip memory use and one on keeping idle chips current.

Editorial take

Claim 1 reaches further than the headline idea suggests. It never says how the data gets shrunk. It only requires that one machine receive an approximated, compressed piece of another machine's work while it is still calculating its own piece, unpack it, and fold it in.

The specific tricks, a matched pair of functions trained together and the LoRA technique, appear only in later claims (4 and 5). So the opening claim, if granted, would reach any compression method that fits that pattern. The same wording repeats for running a model, training one, a computer system, and stored software.

The catch is the word "approximated." The filing admits the results differ slightly from the full-precision answer, and the claims say nothing about how much error is acceptable. Whether that trade pays off is a question for Apple's engineers, and the claims leave them plenty of room.

Get our take in your Top Stories

Liked this breakdown? Add Patentlyze as a preferred source on Google, and our plain-English take shows up more often in your Top Stories the next time Apple patent news breaks.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

17 drawing sheets from US 2026/0311057 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.