Google · Filed Jun 1, 2026 · Published Oct 1, 2026

Google Patents a Two-Stage Image Cleanup Step for Its Vision AI

Google has filed a patent for a neural network that scrubs image data not once but twice before any real analysis begins, a small architectural tweak aimed at making AI vision models more consistent and accurate from the very first step.

A neural network processes an input image through stem and transformer blocks to produce a machine vision result. Drawing from patent filing US 2026/0301398 A1.
A neural network processes an input image through stem and transformer blocks to produce a machine vision result.
See all 5 drawings from this filing ↓
Publication number US 2026/0301398 A1
Applicant Google LLC
Filing date Jun 1, 2026
Publication date Oct 1, 2026
Inventors Manoj Kumar Sivaraj, Neil Matthew Tinmouth Houlsby, Mostafa Dehghani
US classification 382/155
Status when we published Waiting for an examiner (Jun 27, 2026)
Parent application is a Continuation of 18419170 (filed 2024-01-22)
Document 20 claims

What Google's double-normalization vision AI actually does

Every time an AI model looks at a photo, say to identify a face, read a street sign, or flag a defect on a factory line, it starts by chopping the image into small tiles and converting those tiles into numbers. That conversion step is where errors creep in, and those errors compound as the data moves deeper into the system.

Google's patent covers a neural network design that inserts a cleanup pass at two separate points in that early preparation stage. The first pass tidies up the raw image tiles before they get converted into numbers. The second pass tidies up the numbers themselves. The idea is that by the time the data reaches the part of the network doing the actual thinking, it's been stabilized twice, reducing the odds that quirks in the original image throw the whole analysis off.

This isn't a new product you can download. It's an architectural blueprint describing how to wire together the early layers of a vision AI system. The benefit, if it holds up in practice, would show up as more reliable accuracy across varied lighting conditions, image quality, and content types.

From the filing · CLAIM 1
a patch layer that subdivides an input image into a set of image patches; a first normalization layer that generates a set of normalized image patches by performing a first normalization process on each image patch …

Translation: The system chops up the picture and cleans up the pieces right away.

How the two normalization layers prepare image data

The patent describes a neural network built around a stem block, which is essentially the front door through which image data enters the system. That stem block contains four layers arranged in a specific sequence.

  • Patch layer: cuts the input image into a grid of smaller tiles, a standard technique in modern vision transformers (AI models that analyze images by treating each tile like a word in a sentence).
  • First normalization layer: applies a statistical smoothing process (normalization means rescaling values so they sit within a predictable range) to each raw tile before anything else happens.
  • Embedding layer: converts each cleaned tile into a vector embedding, which is a list of numbers that represents the tile's visual content in a form the AI can reason about.
  • Second normalization layer: applies a second round of statistical smoothing to those number-lists before they move on.

After those four steps, the data passes into a transformer block, which is the part of the network that actually performs tasks like classification, detection, or segmentation. The claim is that feeding pre-stabilized data into the transformer reduces training instability and improves final accuracy.

The patent doesn't specify which normalization technique (layer norm, batch norm, and others each have different trade-offs) is used at each stage, leaving that flexible. The key architectural claim is simply that two separate normalization steps, placed precisely where described, produce better results than one.

From the filing · THE ABSTRACT
Each vector embedding of the set of embedding vectors is a projection of a corresponding normalized image patch from the set of normalized image patches onto a visual token.

Translation: Those clean pieces are converted into digital tokens the AI can easily read.

What this means for AI that reads photos and video

For the engineers building vision AI systems at Google, a more stable input pipeline means models that train faster and behave more predictably across different datasets. That has practical knock-on effects for any Google product that relies on image understanding, from Google Photos to Search to cloud APIs that third-party developers use.

For everyday users, the change would be invisible. You'd simply see an AI that makes fewer weird mistakes when the lighting is bad or the photo is slightly blurry. The patent sits at the infrastructure level, the kind of unglamorous-but-load-bearing work that Google keeps filing on in machine vision architecture, where small improvements in data preparation can meaningfully shift accuracy at scale.

Google's 43rd filing in the AI vision work we've tracked since May adds to a run that includes one editing facial expressions and one reading pulse for heart risk.

Editorial take

On the ship-path question, this patent sits about as far upstream from a finished product as possible while still being relevant. No new chips or cameras are required. The entire idea is a software rewiring: arrange four processing steps in a specific order, and the model learns more accurately.

The document describes why the arrangement should work but includes no measurements showing by how much. That gap matters, because a design change with unclear accuracy gains can sit in research indefinitely.

If the improvement proves real in testing, this gets folded into future image-recognition software the way small architectural choices usually do. If not, it stays theoretical.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

5 drawing sheets from US 2026/0301398 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.