Adobe · Filed Feb 25, 2025 · Published Aug 27, 2026 · verified — real USPTO data

Adobe Patents an AI Image Tool That Composes Foreground and Background From Separate Text Prompts

Getting an AI to place a specific subject into a specific background without the two looking like they belong in different photographs is one of the hardest open problems in image generation. Adobe's new patent describes a two-model approach that keeps the subject's details sharp while still making it look at home in the scene.

A foreground object and a background scene combine through an image processing tool to create a single, unified synthetic image. Drawing from patent filing US 2026/0253262 A1.
A foreground object and a background scene combine through an image processing tool to create a single, unified synthetic image.
See all 20 drawings from this filing ↓
Publication number US 2026/0253262 A1
Applicant ADOBE INC.
Filing date Feb 25, 2025
Publication date Aug 27, 2026
Inventors Yusuf Dalva, Yijun Li, Qing Liu, Nanxuan Zhao, Jianming Zhang, Zhe Lin
CPC classification 345/418
Grant likelihood Medium
Examiner CHEN, YU (Art Unit 2613)
Status Non Final Action Mailed (Jul 22, 2026)
Document 20 claims

How Adobe's two-prompt image compositing works

Today's AI image generators struggle when you want precise control over both a subject and its background separately. You describe everything in one long prompt and hope the AI figures it out, but the subject and background often fight each other for the model's attention.

Adobe's approach gives each part its own prompt. You write one description for the foreground object (say, a ceramic coffee mug) and a separate one for the background scene (a sunlit kitchen counter). Two AI models then work together: one focuses entirely on getting the mug right, and the other tries to blend both prompts into a single scene.

The key step is combining what each model has learned before the final image is drawn. The result is a single composed image where the mug looks like it actually belongs on that counter, rather than being pasted in from somewhere else.

From the filing · CLAIM 1
obtaining a first prompt indicating a foreground object and a second prompt indicating a background scene; generating, using a first image generation model, a foreground attention output based on the first prompt …

Translation: The system takes two text prompts to separately design the main subject and its surrounding environment.

How the two AI models coordinate their attention outputs

The patent describes a pipeline that runs two image-generation models in parallel during what is called the attention phase of image synthesis (the stage where the AI figures out which parts of the prompt should influence which regions of the image).

  • Model 1 (foreground model): receives only the foreground prompt and produces a foreground attention output, a set of internal signals that capture how the subject should look, with no interference from the background description.
  • Model 2 (blending model): receives both prompts together and produces a blended attention output, which represents the subject as it should appear within the described scene.
  • Combination step: the two attention outputs are merged into a combined blended attention output. This gives Model 2 a corrected reference that reinforces the subject's details before it renders the final pixel image.

The final image is then generated by Model 2 using that combined signal. The core insight is that running a subject-only model alongside the blending model acts as a generative prior (a strong internal reference) that stops the background description from distorting how the foreground object looks. The claim covers this two-prompt, two-model, combined-attention structure as a complete method.

From the filing · THE ABSTRACT
A first image generation model is configured to generate a foreground attention output based on the first prompt, wherein the foreground attention output represents the foreground object.

Translation: One part of the AI focuses entirely on creating just the main object based on your first description.

What this means for AI photo editing and compositing

For designers and photographers, manual compositing (cutting a subject out of one photo and placing it into another) is time-consuming work that requires matching lighting, color, and perspective by hand. An AI tool that accepted two separate text descriptions and produced a believable composite would compress that workflow significantly. Adobe already ships AI generation features inside Photoshop and Firefly, so a technique like this fits directly into those products.

The claim as written is broad: it covers any method that uses a separate foreground model to produce an attention output that feeds into a blending model's final render, regardless of what AI architecture sits underneath. That breadth would matter if granted, since it could apply to a wide range of text-to-image compositing systems, and new tech patents in AI image generation from Adobe and its peers are adding up into a dense thicket of overlapping claims around creative-tool automation.

This is the 17th Adobe filing we've tracked since May on controllable AI images, following earlier applications like one generating themed design elements and one for text-driven 3D edits.

Editorial take

Claim 1 covers a specific sequence: take two text descriptions, run them through two separate image-generation models to produce two different outputs, combine those outputs, and then generate a final image. That sequence is broad enough to capture any system that follows roughly that same order of operations, where subject and scene are handled separately before being merged.

The word "combined" is where the claim's reach gets tested. Any competing system that blends the two descriptions earlier, before the models produce their separate outputs, could argue it sits outside this claim entirely.

What Claim 1 would block, if granted, is a product that markets itself as doing exactly this: give us a subject description, give us a scene description, and we will composite them for you using separate generation passes. That is a recognizable feature in consumer image tools today, which makes this filing commercially meaningful.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

20 drawing sheets from US 2026/0253262 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.