Salesforce Patents a Way to Automatically Grade and Kill Bad AI Search Results
AI search tools are powerful but unpredictable. Salesforce has filed a patent for a system that grades AI search results in real time and can switch the whole thing off if the quality drops too low.
How Salesforce's search grading system actually works
Ever searched for a product on a store's website and gotten results that had nothing to do with what you typed? That's the problem this patent is designed to catch before it costs the retailer a sale.
Salesforce's system works like a quality inspector for AI-powered search. It looks at every product catalog, figures out how items are organized into categories, and then checks whether the search results an AI tool returns actually match what someone was looking for. If the results are off, the system assigns a low score. If that score falls below a minimum bar, the system can automatically disable the AI search tool entirely.
The idea is to give businesses a safety net. AI search engines don't always return the same results for the same question, which makes them hard to test with traditional methods. This patent describes a way to measure quality even when the AI is being unpredictable, and to act on it without waiting for a human to notice.
… determining, with the computing device, a query score for the non-deterministic search system based on the attributes and values for the categories that were determined to be responsive to the search query and the categories of the items of the search results …
Translation: The system calculates a performance grade by comparing expected product categories against what the search actually returned.
How the system scores queries and pulls the plug on bad searches
The patent describes a pipeline for evaluating what it calls a non-deterministic search system, meaning an AI-powered search tool that can return different results for the same query each time you ask it. Traditional software testing fails here because you can't just check for a fixed expected output.
Here's how the evaluation pipeline works:
- The system ingests a product catalog and automatically extracts categories (like "running shoes"), attributes (like "color" or "size"), and values (like "blue" or "10").
- When a user submits a search query, the system figures out which categories are relevant to that query.
- It then collects the actual results the AI search tool returned and checks which categories those results belong to.
- It computes a query score by comparing what categories should have appeared against what actually did.
If the query score falls below a set threshold, the system can automatically disable the AI search tool and presumably fall back to a more conventional, rule-based search. This is sometimes called a circuit-breaker pattern in software engineering, where a system shuts down a failing component before it causes wider damage.
The scoring approach leans on the structure of the catalog itself as the ground truth, rather than requiring a human to hand-label thousands of "correct" answers for every possible query.
A catalog including items may be received. Categories, attributes for the categories, and values for the attributes, may be determined from the items in the catalog. A search query that was also input to a non-deterministic search system may be received.
Translation: The software starts by organizing a product inventory and monitoring queries sent to an AI search engine.
What this means for online stores running AI-powered search
For any business running an AI-powered product search, bad results aren't just annoying for customers. They translate directly into lost sales, returned orders, and eroded trust. The harder problem is that AI search failures are often invisible until enough customers have already bounced. A system that catches and responds to quality drops automatically could prevent that gap between when things go wrong and when someone notices.
This is particularly relevant as more e-commerce platforms, including those built on Salesforce's Commerce Cloud, swap in AI-based search tools. the pattern in Salesforce's commerce and AI search filings suggests the company is building out infrastructure to manage the reliability of those tools, not just deploy them. That's a more mature approach than simply plugging in an AI model and hoping it works.
Salesforce's 18th filing we've tracked since May in the AI guardrails race follows its work on grading AI answer quality and agent permission controls.
AI-powered product search that misfires costs online retailers in two ways: shoppers leave empty-handed, and the store never knows exactly why. Because AI systems can give different answers to the same question on different days, standard quality checks that expect consistent behavior tend to miss the failures entirely.
Salesforce's approach uses the store's own product catalog as the judge. When search results drift away from the categories of products that should have matched a query, the system flags the gap and can shut the search tool down automatically, without waiting for a human to notice.
The automatic shutdown is the part that deserves scrutiny. A well-calibrated cutoff protects shoppers from a broken experience; a poorly set one could disable a working search engine over a single odd result. How that threshold gets set will determine whether this solves the problem or creates a new one.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
7 drawing sheets from US 2026/0300413 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in