Google Patents a Way to Edit Photos Using Plain-English Instructions
Google has filed a patent for an AI that lets you edit photos by typing what you want changed, no Photoshop skills required. You describe what to target and how to change it, and the system does the rest.
What Google's text-driven photo editing actually does
Every time you want to fix something in a photo, you open a complicated app, search for the right tool, and hope you don't accidentally ruin the whole image. Google's new patent describes an AI system built to skip all of that.
The idea is straightforward: you type two things. First, you describe what in the photo you want to change (say, "the red jacket the person on the left is wearing"). Second, you describe how you want it changed ("make it blue and add a hood"). The AI reads both instructions together, figures out the exact part of the image you mean, and produces an edited version that looks like a real photograph, not a sloppy cut-and-paste.
The patent calls this "referring object manipulation," and Google frames it as a genuinely new kind of problem for AI to solve. Instead of drawing a selection box or using a brush tool, you just talk to the software. That's the whole pitch.
… a machine-learned image manipulation model configured to receive and process an input image and a natural language instruction to generate an edited image in accordance with the natural language instruction …
Translation: An AI tool takes a photo and a text command to produce a modified picture.
How the model reads a two-part instruction to change an image
The system at the center of this patent is a machine-learned image manipulation model, which is an AI trained on enough images and text pairs that it has learned to associate visual regions with language descriptions.
When you give the system an image and a natural-language instruction, that instruction is split into two parts:
- Reference portion: the part that identifies a specific object or region in the photo ("the tall lamp in the corner")
- Target portion: the part that describes what should happen to it ("replace it with a floor plant")
The model processes both the image and the two-part instruction together. It locates the region that matches the reference text, applies the described change, and outputs an edited image. Critically, the patent specifies the result should be photo-realistic, meaning the surrounding context (lighting, shadows, perspective) is preserved so the edit doesn't look artificial.
The patent does not detail the model's internal architecture, but the claim covers any system that takes this two-part natural-language input and produces a contextually appropriate edited image. That makes the claim fairly broad in scope.
… an entirely new problem space of referring object manipulation (ROM). In ROM, a computer system aims to generate photo-realistic image edits regarding two textual descriptions …
Translation: The system uses two text descriptions to make realistic changes to a specific part of a photo.
What this means for everyday photo and image editing
For ordinary users, the practical upside is big. Right now, making a precise change to one object in a photo, without disturbing anything around it, typically requires either professional software skills or a lot of trial and error with consumer AI tools. A system like this would let you make surgical edits by just describing what you want.
For Google specifically, this kind of capability slots naturally into products like Google Photos or the editing tools inside Pixel phones. Google's run of generative image-editing filings signals that the company is building toward a future where the gap between "knowing what you want" and "being able to do it" collapses entirely for photo editing. Whether this specific patent shapes a shipping product is another question, but the direction is clear.
Google files its 58th application we've tracked since May in our photo and video editing watchlist, building on earlier ideas like rebuilding video from sample frames and distance-based photo sharpening.
Claim 1 covers any system that takes an image, a phrase identifying something in it, and a phrase describing what to change, then produces a photo-realistic result. It does not require a specific underlying model, a specific way of finding the object, or a specific training approach. That is a wide net.
In practice, the two-part instruction structure (point at something, say what to do with it) is the most natural way any product would accept an editing command, so the claim reaches across a large share of conversational image editing tools, not just one technical implementation.
The only real constraint is photo-realism, and that bar has been falling fast as the underlying technology matures. If this claim is granted as written, it creates substantial commercial territory around one of the most intuitive interactions a user could reach for.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
11 drawing sheets from US 2026/0289845 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in