Zoox Patents a Way to Steer Autonomous Vehicles With Hand Gestures From Afar
Instead of a keyboard or joystick, a remote operator could point at a 3D map of a self-driving vehicle's surroundings and direct it with hand gestures. That's the core idea in Zoox's latest patent filing.
What Zoox's remote gesture control actually does
Remote operators who oversee self-driving vehicles today typically work with flat screens, clicking or typing commands the way you'd manage a video call. Zoox wants to change that by giving operators a wearable device that shows a live, three-dimensional picture of wherever the vehicle is, so they can use natural hand movements to give it instructions.
The idea is straightforward: the vehicle's sensors feed data back to the operator in real time. That data gets turned into a 3D map the operator can see and interact with. Spot a pedestrian about to step into the road? Point at them. The system reads your gesture, figures out what command it maps to, and sends that instruction to the vehicle.
This is aimed at the human safety net behind autonomous vehicles, the people who step in when the car needs help. Rather than hunting through menus, they'd interact the way you'd gesture at a table in front of you.
… receiving, second natural user input data of the user, wherein the second natural user input data includes an indication of a feature in the representation of the environment; determining, based at least in part on the indication of the feature in the representation, an action associated with the vehicle; and transmitting information to the vehicle to perform the action.
Translation: The system watches a hand gesture pointing at something in the virtual scene and tells the car to do a task based on that.
How the system maps gestures to vehicle commands
The patent describes a two-stage interaction loop built around a wearable computing device (think AR headset or smart glasses worn by a human operator sitting remotely).
First, the vehicle sends back sensor data from its cameras, lidar, or radar. That data is combined into a 3D representation of the environment around the vehicle, including detected people and objects, and displayed on the operator's wearable screen.
The system then tracks what Zoox calls natural user input, meaning body movements like hand gestures, head turns, or gaze direction, rather than button presses. The patent breaks this into two layers:
- First input: managing the 3D view itself (panning, zooming, rotating to look around the scene).
- Second input: pointing at or selecting a specific feature in the scene, such as a pedestrian or an intersection, to trigger a vehicle action.
Once the system identifies which object or location the operator is indicating, it determines the appropriate vehicle command and transmits it. The patent does not specify the exact gesture vocabulary, leaving room for the underlying gesture-recognition engine to define that.
The representation may include features of the environment, such as people and/or objects. The representation may be displayed at a user interface. In some instances, the user interface may be associated with a wearable computing device and may be associated with a user, such as a remote operator.
Translation: A remote worker wearing a headset views a 3D simulation of what is happening around the autonomous vehicle.
What this means for human oversight of self-driving cars
As autonomous vehicles edge toward commercial deployment, the industry has largely settled on keeping humans in the loop for edge cases. How those humans actually communicate with a vehicle in a hurry matters enormously. A clunky interface during a split-second decision is a safety risk, not just an inconvenience. If a gesture-based system genuinely reduces the steps between "operator sees a problem" and "vehicle receives a command," that gap reduction has real consequences for safety.
For you as a passenger or bystander, this is about whether the human backstop in an autonomous system can respond fast enough and clearly enough to matter. Zoox's investment in remote-operation interfaces suggests the company sees human oversight as a long-term feature, not a transitional crutch while full autonomy catches up.
This is the 22nd Amazon filing we've tracked since May on our self-driving sensing race, which already includes Zoox applications like one that maps roads live and one for dodging collisions.
Waving your hand to steer a remote vehicle is faster than clicking through menus, and speed matters when a vehicle is moving through traffic and decisions can't wait. But the patent never explains how the system tells a real instruction from an accidental movement, and that gap is the central unresolved cost of the whole design.
That cost is not abstract. A misread gesture sent to a moving vehicle is a serious failure, not a minor one, so the speed advantage only holds if the system gets the command right on the first try.
The spatial interface concept is strong enough to be worth building, and the problem it solves, controlling a vehicle through a three-dimensional view with your hands, is real. Whether the tradeoff reads as worth it depends entirely on reliability the patent does not yet address.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
6 drawing sheets from US 2026/0267329 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →