Apple Patents a Way to Track Objects in 3D Even When Depth Data Goes Missing
Depth sensors don't always fire on every frame, and when they miss, most 3D tracking systems fall apart. Apple's new patent describes a way to keep the estimate going anyway, using what the camera already knows.
How Apple keeps 3D object tracking alive without a depth sensor
Ever tried to measure a piece of furniture with your phone camera, only for the AR overlay to wobble or snap out of place? That's often because your phone's depth sensor didn't catch every frame, leaving the software guessing.
Apple's patent describes a system that handles those gaps more gracefully. When the depth sensor does fire, the system locks in a solid estimate of how big an object is and where it sits in 3D space. When depth data is missing in a later frame, the system leans on that previous size estimate and the regular camera image to keep tracking the object's position. No gap, no wobble.
The practical effect is that 3D measurements and AR overlays stay stable even as your device moves around a room and the sensors skip frames. For everyday tasks like measuring a couch before you buy it or placing a virtual object in your living room, this kind of consistency is what makes the experience feel reliable rather than frustrating.
For each frame of the sequence of frames that includes image data and excludes depth data, the size estimate is determined based on a prior size estimate from a prior frame, and the 3D position is determined based on the size estimate of the representation of the object and a portion of the image data.
Translation: The system keeps track of objects in 3D by relying on historical size data whenever the depth sensors stop working for a specific frame.
How the system fills gaps when depth data drops out
The patent describes a frame-by-frame pipeline for estimating an object's 3D position and physical size using a mix of camera images and depth data.
When a frame includes both image data and depth data (from something like a LiDAR or time-of-flight sensor), the system does the full calculation: it uses the device's known orientation and position in space (pose data, essentially where the phone is pointing and how it's tilted) together with the depth reading to pin down exactly where the object is, and it uses the depth plus the image to calculate how large the object appears to be in the real world.
When a frame has image data but no depth data, the system does not give up. Instead:
- It pulls the size estimate from the most recent frame that did have depth data (the "prior size estimate").
- It uses that known size as a reference anchor, then reads the object's apparent size in the current camera image to back-calculate how far away the object must be.
- From that inferred distance, it reconstructs the object's 3D position without needing fresh depth sensor input.
The result is a continuous stream of position and size estimates even when the depth sensor misses frames, which happens regularly on consumer devices during fast movement or in low-contrast scenes.
What this means for AR on iPhones and Vision Pro
Depth sensors on iPhones and the Vision Pro headset are useful but imperfect. They miss frames, they struggle at certain distances, and they drain power. Any AR or measurement feature that depends on depth alone will stutter when the sensor stumbles. Apple's approach treats depth data as a helpful input when available, not a hard requirement on every single frame, which makes the whole system more resilient in real conditions.
For consumers, this points toward more reliable versions of features like the Measure app or Vision Pro's spatial awareness. For developers building AR experiences, a steadier object-tracking foundation means fewer edge cases to code around. Apple's run of spatial-computing filings shows the company treating 3D object understanding as infrastructure, not a feature.
Apple's 75th filing we've tracked in our AR glasses work since May adds to a run that includes one on shifting 3D photo depth and one on varied meditation audio.
The problem here is real and specific. Current depth sensors on phones and headsets simply do not produce clean, continuous data, and every missed frame is a chance for an AR object to jump or a size measurement to drift. That is a frustrating user experience, and it gets worse in motion-heavy situations, exactly when people are most likely to be using a measuring or placement tool.
This patent's answer is proportionate to the problem. Instead of demanding better hardware, it makes smarter use of what the hardware already gives you: lock in a size estimate when you can, then use basic geometry to hold position when you cannot. That is a practical engineering trade-off, not a theoretical one.
The honest caveat is that this only works as well as the initial size estimate. If the first depth reading is off, the downstream frames inherit that error. Still, for the intended use case, a slightly imperfect but stable estimate beats an unpredictable one.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
7 drawing sheets from US 2026/0301215 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in