New Google Patents · Filed Apr 20, 2026 · Published Sep 3, 2026 · verified — real USPTO data

Google Patents a Way to Edit AI-Generated Images by Gesturing Over Them

Editing an AI-generated image today usually means typing a new prompt or hunting through menus. Google wants to let you just wave your finger over the part you want changed.

A user's hands gesture over a tablet displaying an AI-generated image of a UFO, indicating modification. Drawing from patent filing US 2026/0259662 A1.
A user's hands gesture over a tablet displaying an AI-generated image of a UFO, indicating modification.
See all 9 drawings from this filing ↓
Publication number US 2026/0259662 A1
Applicant GOOGLE LLC
Filing date Apr 20, 2026
Publication date Sep 3, 2026
Inventors Ramprasad Sedouram, Siyan Khader Sameema, Ajay Prasad, Karthik Srinivas
CPC classification 345/173
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (May 27, 2026)
Parent application is a Continuation of 18781611 (filed 2024-07-23)
Document 18 claims

What Google's gesture-based AI image editing actually does

Right now, changing part of an AI-generated image means going back to a text box and rewriting your request, or finding the right editing tool buried in a menu. The image itself is just a picture you look at, not something you can poke to reshape it.

Google's patent describes a different approach: you gesture over a specific part of the image on your screen, and the AI figures out what you want to change about that area and changes it. Swipe, circle, or tap a portion of the picture, and the system generates a new version of just that section. No separate editing panel, no re-typing your original prompt.

The system can handle preset gestures with fixed meanings, or it can interpret your motion more freely using a model trained to read what a gesture likely intends. Either way, your finger becomes the editing tool, directed straight at the content.

From the filing · CLAIM 1
… determining, subsequent to the natural language input being received at the client application, that an input interface of the client computing device, or another computing device, has received an input gesture, wherein the user performs the input gesture by motioning over a particular portion of the display interface that is providing a rendering of the image file; …

Translation: The system detects when a user makes a physical hand motion over a specific part of the displayed picture.

How the system reads your gesture and rewrites the image

The patent describes a pipeline that begins with a normal AI image request: you type something in a chat or prompt interface, and a generative model (an image diffusion model, the kind that turns text into pictures) produces an image and displays it.

After the image is on screen, the system monitors the input interface of your device, meaning your touchscreen, stylus, or even a motion sensor on a nearby device. When it detects a gesture directed at a particular region of the displayed image, that gesture becomes a new input to the AI.

The gesture is processed one of two ways:

  • Preset gestures: a circle might mean "make this area simpler," a scratch might mean "remove this element" -- fixed mappings the system recognizes.
  • Model-interpreted gestures: a separate AI model reads the shape, speed, and direction of your motion and infers what edit you probably want, even if it doesn't match a preset.

The system then generates additional generative output, which is a modification instruction or a replacement image patch. The client app swaps out that portion of the original image with the revised version. The interaction is meant to feel direct: touch the thing you want to change, and the AI changes it.

From the filing · THE ABSTRACT
Various different input gestures can be performed by a user to refine a generative output to be simpler, more complex, to include an image, to modify a generated image, and/or otherwise modify the generative output.

Translation: Users can make various hand movements to make the generated image simpler, more complex, or edited in different ways.

What this means for AI image tools on phones and tablets

For anyone using AI image tools on a phone or tablet, this is the kind of friction that currently makes editing feel like a chore. Typing a follow-up prompt that targets one corner of an image is awkward, and most mobile AI interfaces aren't built for surgical text-based edits. A gesture-first approach fits how people naturally use touchscreens.

Google's run of generative-AI interface filings suggests the company is thinking hard about how AI tools behave after the first output, not just how they generate it. The more interesting question is whether gesture interpretation is reliable enough to feel useful in practice. If the AI misreads your circle as a scratch, you get a wrong edit instead of a wrong answer, which can be harder to undo.

Google's 45th filing we've tracked since May in the AI photo editing race follows one preserving detail on weak hardware and one rebuilding video from motion layers.

Editorial take

The core payoff here is real and specific: you would notice this during the "almost right" moment, when an AI image is 90% what you wanted but one element is off. Right now that moment sends you back to the keyboard. A gesture shortcut, if it works, keeps you in the image.

The harder problem the patent acknowledges is ambiguity. Gestures don't carry inherent meaning the way words do, so the system needs either a rigid dictionary of movements or a model smart enough to infer intent from a swipe. The patent hedges by proposing both, which is honest but also signals that neither is a solved problem.

For touch-first devices like tablets and phones, this is a genuinely practical direction. The gap between "good enough to ship" and "reliable enough that users trust it" will determine whether this feels like a feature or a frustration.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

9 drawing sheets from US 2026/0259662 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.