Google Patents a Two-Model AI System for Ranking Search Results
Google has filed a patent for a search ranking method that runs your query and every candidate document through two separate AI models, then feeds the comparison into a third neural network to decide what you see first.
How Google's dual-AI search ranking actually works
A researcher types a question into Google. The engine comes back with ten blue links in under a second. What decides which link lands at the top? That ordering is the whole game, and Google is constantly rebuilding the machinery behind it.
This patent describes a method where your query is processed by one AI model and each candidate document is processed by a second, separate AI model. The system then compares the outputs, building a kind of grid that captures how closely your words and each document's words line up. A third neural network reads that grid and assigns every document a relevance score. The highest scores rise to the top of your results.
The key idea is that the query side and the document side each get their own specialized model rather than a single shared one. That separation lets each model focus on what it does best before the comparison happens.
… computing, by the computing system, a similarity matrix based on a first content associated with the query generated by a first transformer and a second content associated with a plurality of documents generated by a second transformer …
Translation: Two separate AI transformers process the search query and documents to compare how similar they are.
How the similarity matrix drives document scoring
The patent describes a four-step pipeline for ranking search results.
- Step 1, Query encoding: A "first transformer" (a type of AI model that reads text by paying attention to word relationships across an entire sentence at once) processes the incoming search query and produces a mathematical representation of its meaning.
- Step 2, Document encoding: A "second transformer," separate from the first, processes each candidate document in the same way, producing its own set of mathematical representations.
- Step 3, Similarity matrix: The system computes a similarity matrix (think of a spreadsheet where each cell records how closely one piece of the query matches one piece of a document). This matrix captures fine-grained, token-level overlaps rather than just a single summary number.
- Step 4, Neural network scoring: A separate neural network reads the full matrix and outputs a relevance value for each document. Those values determine the final ranking order.
Using two distinct transformers rather than one shared model is the structural choice the patent centers on. The query encoder and document encoder can be trained or tuned independently, which in theory lets each specialize. The similarity matrix then acts as a bridge, carrying all the detailed comparison data into the scoring network rather than collapsing it into a single number too early.
… processing, by the computing system, the similarity matrix via a neural network to generate relevance values respectively corresponding to the plurality of documents …
Translation: A neural network uses those comparison results to score how well each document answers the search query.
What this means for search quality and competition
Search ranking is one of the most commercially sensitive problems in software. Small changes in what appears at position one versus position three can shift enormous amounts of web traffic and ad revenue. A method that extracts more signal from the relationship between a query and a document before making a ranking call could produce meaningfully better results for users, especially on long or ambiguous queries where a single summary score misses nuance.
For you as a searcher, the practical promise is fewer situations where the top result technically contains your words but doesn't actually answer your question. Google's track record in search-ranking patents suggests this is incremental refinement of an existing system rather than a wholesale replacement, but incremental at Google's scale affects billions of queries a day.
Google's 25th filing we've tracked since May on AI agents working together follows earlier work on agents watching software connections and routing questions to the right model.
Running two separate language models instead of one means roughly twice the computing cost before the ranking even begins, and search engines have to return answers in fractions of a second. That is a real bill, and this document does not make the case that the accuracy improvement is large enough to cover it.
The subtler risk is coordination. Two models that each learn language separately have to stay compatible with each other, and whenever one gets updated the other may drift out of alignment. That kind of invisible mismatch is exactly where systems like this tend to fail in daily use rather than breaking loudly in testing.
For a company running one of the world's largest computing infrastructures, the trade is at least defensible. For almost anyone else trying to build on this approach, the overhead is the product.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
10 drawing sheets from US 2026/0300306 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in