Apple · Filed Jan 20, 2026 · Published Aug 6, 2026 · verified — real USPTO data

Apple Patents a Camera That Reframes Shots Using Gestures and Voice Together

Say 'closer' while waving your hand and the camera zooms in. Say 'wider' with the same wave and it zooms out. Apple is patenting a camera interface where the same gesture does different things depending on what you say at the same time.

Apple Patent: Gesture Plus Voice Camera Reframing — figure from US 2026/0230700 A1
Figure from the official USPTO publication.
See all 40 drawings from this filing ↓
Publication number US 2026/0230700 A1
Applicant Apple Inc.
Filing date Jan 20, 2026
Publication date Aug 6, 2026
Inventors Fiona P. O'LEARY, Sean Z. AMADIO, Jeffrey T. BERNSTEIN, Mark K. HAUENSTEIN
CPC classification 348/222.1
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Feb 20, 2026)
Parent application Claims priority from a provisional application 63755169 (filed 2025-02-06)
Document 22 claims

How Apple's gesture-plus-voice camera framing works

Imagine trying to take a group photo without a tripod. You hold up your hand to frame the shot, but right now your phone can only guess what you want. Apple's new patent describes a system that pairs your hand gesture with whatever you're saying out loud, so the camera knows exactly how to reframe.

The key idea is that the same gesture can trigger different results depending on your spoken words. Wave your hand and say one thing, and the camera adjusts one way. Wave your hand and say something different, and it adjusts in an entirely different direction. Your words and your movement work together as a combined command.

The patent also covers a few related tricks: automatically repositioning something in the frame when the camera spots it, and letting you undo a framing change by making an "air gesture" (a hand movement in open space, not touching any screen). Together, these suggest Apple is thinking seriously about hands-free, voice-assisted camera control.

How spoken words change what a gesture actually does

The core of this patent is a two-channel input system for camera framing. The camera viewfinder watches for hand gestures using the device's cameras, while the microphone listens for speech at the same time. Crucially, the system evaluates both signals together to decide which framing action to take.

How the logic works:

  • If the system detects a gesture accompanied by a first specific speech input, it shifts the camera view to a second field-of-view (a new angle or zoom level).
  • If the system detects the same gesture but accompanied by a different speech input, it shifts the camera to a third, distinct field-of-view.
  • This means one gesture can map to multiple outcomes, disambiguated entirely by what you say.

The patent also describes two additional behaviors. First, an automatic content-repositioning feature: if the camera notices a particular object or subject already in the frame, it can automatically shift the view to place that content more deliberately. Second, an "air gesture" undo function (a hand movement made in space without touching any surface) that reverses recent framing changes.

The system is described as running on a computer system connected to display generation components and cameras, which in Apple's language typically covers iPhone, iPad, and spatial computing devices like Vision Pro.

We find one patent like this every day. Get the best of each week in your inbox, free →

What this means for hands-free photography on Apple devices

For most people, reframing a photo or video shot mid-capture means physically moving the device or tapping a screen, both of which can blur the image or ruin the moment. A system that reads your hand movements and your words together could make hands-free camera control genuinely reliable, rather than a party trick that only works half the time.

This matters most in contexts where your hands are busy or your phone is mounted at a distance: recording yourself, capturing a family moment, or shooting video solo. Apple already has gesture and voice recognition across its product line. Combining them into a single camera-control layer is a natural next step, and this patent suggests the engineering work to do it is underway.

Editorial take

This is a genuinely thoughtful approach to a real problem: gesture-only camera control is too ambiguous, and voice-only control can feel awkward in public. Fusing both signals at the decision point is clever, and the undo-via-air-gesture detail shows the team has thought about what happens when the system gets it wrong. Whether this ships as a Camera app feature or something specific to Vision Pro spatial photography is the interesting open question.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

40 drawing sheets from US 2026/0230700 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.

Editorial commentary on a publicly published patent application. Not legal advice.