Sony Patents a Way to Stream 3D Object Scans Without Re-Sending Unchanged Data
Streaming a full 3D scan of an object every single frame is expensive. Sony's new patent cuts that cost by flagging the parts of the scene that haven't moved and telling the player it can skip re-downloading them.
What Sony's 3D streaming shortcut actually does
Imagine watching a 3D scan of a person walking through a room: the furniture barely moves, but the person moves constantly. Today, streaming systems often re-send all that 3D data, including the couch that hasn't budged, on every update. That wastes bandwidth and makes playback harder.
Sony's patent describes a system that labels each chunk of a 3D scene as either static (not changing) or dynamic (changing). A small "change report" file travels with the 3D data, telling the receiving device which chunks it can reuse from before and which ones it actually needs to download fresh. The player only does heavy decoding work on the parts that actually changed.
The result is a lighter workload on the device playing back the 3D content. You'd potentially get smoother, faster 3D playback without needing a dramatically more powerful processor.
… generate a file storing a bit stream of encoded data acquired by encoding the point cloud, the first information, and the second information; and transmit the bit stream to a receiving device …
Translation: It bundles the 3D data and change details into a single file and sends it to the viewer.
How Sony's encoder tracks static versus moving regions
The patent centers on point clouds (three-dimensional scans where an object or scene is represented as millions of tiny dots in space). Streaming point clouds is data-heavy, because each dot has a position and often color information.
Sony's system divides the 3D space into discrete tiles (independently decodable chunks of the scene, think of them like puzzle pieces covering the 3D volume). The encoder then generates two pieces of metadata alongside the actual compressed scan data:
- First information: a flag per tile region indicating whether its relationship to the full point cloud is static (unchanged frame to frame) or dynamic (shifted or altered).
- Second information: structural details about the tile regions, but only for the dynamic ones. For static regions, the system simply records how many there are, without re-sending the full map of where they are.
The receiving device reads the first-information flag first. If a region is static, it knows it can reuse previously decoded tile data without re-extracting or re-decoding anything. If a region is dynamic, it pulls the fresh tile data and decodes only that. All of this is packaged in a single file alongside the encoded bitstream, so no separate channel is needed.
… in a case in which the relationship described above is static, data of tiles composing the three-dimensional spatial regions constructing a point cloud is extracted on the basis of the second information, and the data is decoded.
Translation: When parts of the 3D object stay the same, the system reuses old data instead of downloading it again.
What this means for 3D video and spatial computing
Point cloud streaming sits at the foundation of technologies like volumetric video (true 3D video of real people and objects) and spatial computing experiences. Right now, the processing load of decoding full 3D streams limits either the quality or the hardware tier required to play them back. A system that skips redundant work on static regions directly shrinks that load.
Sony's steady investment in point-cloud compression suggests the company is building toward practical volumetric content pipelines, likely for entertainment and communication. For you as a viewer or user, the practical promise is richer 3D experiences on devices that don't require industrial-grade computing power to run them.
That makes this the tenth sensor patent from Sony we've tracked since June, a group that includes one on a self-cleaning gas sensor and one on an adaptive electromagnetic scanner.
Streaming a 3D scan of a real room takes enormous computing power, and most of that work gets repeated even when large parts of the scene haven't moved. Sony's patent describes a way to label which regions of a scene are stable over time so the playback system can skip redundant work automatically.
No new hardware is required. The approach lives in encoding logic and metadata, meaning it could be layered onto existing streaming pipelines. Sony already has deep involvement in the standards bodies where this kind of metadata would need to land, which shortens the path considerably.
A video call where one person moves through an otherwise static room is a natural first target: common, practical, and well-matched to exactly what this optimization handles best.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
42 drawing sheets from US 2026/0268531 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →