Meta Patents a Way to Tweak Your Avatar's Look Using Text Descriptions
Meta has filed a patent that uses AI image editing, guided by text prompts, to modify photos of a person before those photos are converted into a 3D avatar, letting the system change things like hair or skin appearance while leaving actual facial features untouched.
How Meta adjusts avatar photos before building your 3D face
You're setting up a virtual reality profile and you want your avatar to look like you, but with different hair or a cleaner background. Right now, getting that right usually means manual editing or accepting whatever the system generates.
Meta's patent describes a process where AI edits your reference photos first, using text instructions like "shorter hair" or "no glasses," before the system builds your 3D avatar. Crucially, it locks down the parts of the photo that are actually you, things like your eyes, nose, and mouth, so those don't change during the edit.
The result is a 3D model built from photos that have been styled the way you want, without distorting your actual face. The system works across a whole set of photos at once, keeping the adjustments consistent so your avatar doesn't look different from every angle.
… the first segmentation mask identifying a changeable portion and a non-changeable portion of the first two-dimensional image, the non-changeable portion comprising the first facial feature …
Translation: The system figures out what part of the face to edit while keeping key features untouched.
How the segmentation mask protects your real features
The patent describes a three-step pipeline for building a personalized 3D avatar from a set of ordinary photos.
Step one: masking. The system scans a photo and identifies specific facial features, things like the shape of your eyes or the contour of your face. It then creates a segmentation mask (essentially a stencil that marks which parts of the image are off-limits). Those protected areas, called the "non-changeable portion," won't be touched by the AI editor.
Step two: editing with text prompts. An image modification model (a generative AI similar to tools like Stable Diffusion) takes the photo, the mask, and two text instructions: a positive prompt (what to add or emphasize, such as "curly hair") and a negative prompt (what to remove or avoid, such as "glasses"). It edits only the unprotected parts of the image, then replaces the original photo in the set with the edited version.
Step three: 3D reconstruction. A three-dimensional modelling model (software that infers depth and geometry from multiple 2D photos taken at different angles) takes the full updated set of images, now consistently styled, and builds a 3D model of the person.
- Facial features are preserved across edits
- Text prompts control the style changes
- The 3D build uses the edited photos, not the originals
… adjusting, using an image modification model, from a first numerical representation of the first two-dimensional image, the first segmentation mask, a positive prompt, and a negative prompt, the first two-dimensional image …
Translation: AI uses text instructions to alter the photo based on what you want added or removed.
What this means for Meta's avatars in VR and social apps
For anyone using Meta's VR or social platforms, this describes a path toward avatars that look more like you and less like a generic approximation. Instead of manually sculpting a digital face from sliders and menus, you could describe what you want in plain words and have the system handle the photo prep automatically.
Meta's steady investment in avatar personalization sits inside a broader push to make VR social spaces feel less like video games and more like real interactions. If the 3D avatar pipeline described here can reliably separate "things about me" from "things I'd like to change," it solves one of the more persistent frustrations with current avatar tools: the system either copies you too literally or drifts too far from your actual face.
This is the second Meta filing we've tracked in the AI photo editing race since July, following their 3D object placement application.
Claim 1 is structured broadly. It covers any computer-implemented method that (1) generates a segmentation mask from a facial feature, (2) uses a text-prompted image model to edit only the unprotected region, and (3) feeds the edited image set into a 3D modelling model. There is no claim language tying this to a specific AI architecture, a specific platform, or even a specific type of photo. That breadth means the claim, if granted as written, could cover a wide range of AI-assisted avatar workflows, not just Meta's own implementation.
What it would practically block is any pipeline where a competitor uses text-guided AI editing as a pre-processing step before 3D avatar generation, with facial features held fixed by a mask. That is an increasingly common pattern as generative AI tools get woven into avatar creation software. The claim's specificity lives in the combination of mask, positive prompt, negative prompt, and 3D reconstruction together, which is a narrower combination than it first appears, but still broad enough to matter.
The honest read is that this is a defensive filing in a space Meta clearly cares about. The underlying technology, AI image editing plus 3D reconstruction, is not new on its own. The claim earns its scope from the specific arrangement of those pieces, and whether that arrangement is obvious over existing tools will be the central question for an examiner.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
6 drawing sheets from US 2026/0268595 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in