Microsoft · Filed Mar 31, 2025 · Published Oct 1, 2026 · verified — real USPTO data

Microsoft Patents a Frequency-Aware Training Method for AI Image Generators

AI image generators learn by practicing how to clean up noisy pictures, but they typically treat all parts of an image the same way when adding that noise. Microsoft's new patent describes a different approach: add noise according to how each frequency band in the image actually behaves, which could make the learning process more accurate from the start.

Microsoft Patent: Frequency-Domain Diffusion Model Training — figure from US 2026/0300715 A1
Figure from the official USPTO publication.
See all 4 drawings from this filing ↓
Publication number US 2026/0300715 A1
Applicant Microsoft Technology Licensing, LLC
Filing date Mar 31, 2025
Publication date Oct 1, 2026
Inventors Fabian Maximilian FALCK, Sushrut KARMALKAR, Richard Eric TURNER, Javier ZAZO RUIZ, Edward William Frederick MEEDS, Kiarash ZAHIRNIA, Rachel Shannon LAWRENCE, Teodora Plamenova PANDEVA
CPC classification 706/25
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (May 9, 2025)
Document 20 claims

What Microsoft's frequency-domain AI training actually does

You're scrolling through AI-generated images, and something looks subtly off about the fine details. The textures are muddy, the edges slightly wrong, even though the overall composition looks fine. That gap between broad strokes and fine details is partly a training problem.

Here's the issue: AI tools like image generators (Midjourney, DALL-E, Stable Diffusion-style systems) learn by practicing on corrupted versions of real images, then figuring out how to restore them. The corruption step, called adding noise, has traditionally treated every part of an image the same. But an image's sharp edges and its soft gradients behave very differently. Microsoft's patent proposes adding noise in a way that respects those differences, targeting each frequency band based on how much natural variation that band actually carries.

The idea is that if you train an AI on more realistic, structured noise, it should learn to reconstruct images more faithfully. Think of it less as scrambling a puzzle randomly and more as applying the right amount of blur to each type of piece before asking the AI to reassemble it.

From the filing · CLAIM 1
… applying: a first frequency-specific degradation to a first frequency band of the frequency-domain representation based on a first variance of the first frequency band, and a second frequency-specific degradation to a second frequency band of the frequency-domain representation based on a second variance of the second frequency band …

Translation: The system adds targeted noise to different frequency layers depending on how much they vary.

How the patent splits images by frequency before adding noise

The patent describes a new training procedure for diffusion models (the family of AI systems behind most modern image generators). Diffusion models learn in two phases: a forward process that progressively corrupts a clean image with noise, and a reverse process that learns to undo that corruption step by step.

Microsoft's contribution is to move both processes into the frequency domain (meaning the system first converts an image into a mathematical representation that separates slow, sweeping changes from rapid fine-grained ones, similar to how audio equalizers separate bass from treble). Instead of adding uniform noise, the system measures the variance per frequency band (how much that part of the image naturally fluctuates across training data) and scales the noise accordingly.

The forward process degrades the frequency-domain version of each training image using those band-specific variance measurements. The reverse process starts from a noisy frequency-domain signal, also scaled to match those variances, and iteratively refines it toward a clean output.

The key claim is that both directions use the same underlying principle: frequency-aware, variance-based scaling. This keeps the noise at each stage statistically consistent with how real image data is distributed, rather than applying a one-size-fits-all corruption schedule.

From the filing · THE ABSTRACT
… a forward process systematically introduces noise by transforming each training sample into a frequency representation and injecting frequency-specific perturbations based on a measured variance per band.

Translation: Training data gets scrambled across different frequency levels to help the AI learn.

What this means for AI-generated image quality

For you as a user of AI image tools, the direct promise is better fine detail. If the training noise more accurately reflects how image data is structured, the model should learn a cleaner mapping from noise to image, which could mean sharper textures, more accurate edges, and fewer artifacts in generated outputs, without necessarily needing more compute or more training data.

Microsoft's long bet on generative AI infrastructure shows up clearly here. This is a low-level training technique, not a user-facing feature, so any payoff would come embedded inside future models rather than as a visible update. The improvement, if it holds up in practice, would be the kind of thing you notice only by comparison: a model trained this way versus one trained the old way, side by side.

Microsoft's fifth filing we've tracked since July in the AI photo editing race connects to one that reads photo content and one targeting exact color in AI image generators.

Editorial take

The clearest place anyone would notice this improvement is in the small details AI image generators routinely fumble: individual threads in a knit sweater, grain in a wooden table, hair separating into distinct strands. This training approach handles those fine textures more carefully, treating different levels of detail with different amounts of controlled noise during learning rather than applying the same blunt treatment across the board.

Whether that translates into a visible difference depends on how well the technique survives the journey from research to shipping product. Foundational training improvements are invisible until they show up as fewer disappointed prompts and less manual touch-up work.

If it works as described, the person who benefits most is anyone who has given up on AI-generated images for anything requiring fine texture and gone back to doing it by hand.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

4 drawing sheets from US 2026/0300715 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.