Nvidia · Filed Feb 5, 2025 · Published Aug 6, 2026 · verified — real USPTO data

Nvidia Patents an AI That Drops Objects Into Photos Without Manual Editing

Adding an object to a photo usually means carefully painting a selection mask by hand. Nvidia's latest patent describes a system that figures out where the object should go, generates it from a text description, and drops it in automatically.

Nvidia Patent: AI Object Insertion Without Manual Masks — figure from US 2026/0228932 A1
Figure from the official USPTO publication.
See all 18 drawings from this filing ↓
Publication number US 2026/0228932 A1
Applicant Nvidia Corporation
Filing date Feb 5, 2025
Publication date Aug 6, 2026
Inventors Raju Tarachand Wagwani, Pankaj Ratnakar Kadtan
CPC classification 345/418
Grant likelihood Medium
Examiner AMIN, JWALANT B (Art Unit 2612)
Status Docketed New Case - Ready for Examination (Mar 14, 2025)
Document 20 claims

How Nvidia's auto-inpainting skips the manual masking step

Imagine you have a photo of a living room and you want to see what a particular lamp would look like sitting on the side table. Today, doing that convincingly in a photo editor means manually tracing the area where you want the object to appear. It's tedious, and one wrong stroke makes the result look fake.

Nvidia's patent describes an AI pipeline that skips that step entirely. You type a description of the object you want (say, "a modern floor lamp"), and the system figures out on its own where in the photo that object would realistically belong. It then generates a version of the object, resizes and recolors it to match the scene, and drops it in.

You can get back several different versions showing the object in slightly different positions, and you pick the one that looks best. The goal is to make photo compositing feel less like surgery and more like a prompt.

How the segmentation and text-to-image pipeline fits objects in

The system works in two main stages, both driven by machine learning models.

First, a segmentation model (a type of AI that categorizes every part of an image, like labeling one region as "floor," another as "wall," and another as "table surface") scans the input photo and maps out what kinds of surfaces and spaces are available. That map is then used to figure out plausible spots for a new object.

Second, a text-to-image model (the same family of AI that powers tools like DALL-E or Stable Diffusion) generates candidate images of the requested object. Critically, it does this with awareness of the scene context, so the generated object already accounts for the available regions. The segmentation model then runs again on those generated images to cleanly extract just the object itself, cutting away the background.

After the object is placed, the system applies post-processing adjustments:

  • Scaling the object to match the perspective of the scene
  • Adjusting aspect ratio to fit naturally into the available space
  • Matching the color profile of the photo so lighting and tone look consistent

The output is multiple variations of the final image, giving the user options to choose from.

We find one patent like this every day. Get the best of each week in your inbox, free →

What this means for photo editors and AI image tools

Right now, high-quality object compositing in photos is a skill that takes years to develop in tools like Photoshop. This patent points toward a future where you describe what you want in plain text and an AI handles the placement, sizing, and color matching automatically. That's a meaningful shift for anyone who does product photography, interior design visualization, or social media content without a professional editing background.

For Nvidia, this fits squarely into its push to build AI-powered creative tools. The company already sells infrastructure that runs image-generation models at scale, and owning the patents around context-aware placement could matter as those tools move from the cloud into consumer-facing applications. If you use any AI image editor in the next few years, there's a real chance the "add this to my photo" button is doing something close to what this patent describes.

Editorial take

This is a genuinely practical patent in a space that's currently frustrating for non-experts. The manual masking step is one of the biggest friction points in photo compositing, and automating it well would make AI image tools substantially more useful. Whether Nvidia ships this directly or licenses the approach, it's worth tracking.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

18 drawing sheets from US 2026/0228932 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.

Editorial commentary on a publicly published patent application. Not legal advice.