Salesforce Patents a Scoring System for AI Agents Handling Customer Service Tasks
Before you trust an AI agent to handle your customer support queue, someone has to figure out whether it's actually any good at the job. Salesforce is patenting a formal scoring system to do exactly that.
What Salesforce's AI agent grading system actually does
Imagine you're a manager who just hired ten temp workers through different agencies, and you need a fair way to grade each one before deciding who stays. You give them all the same test tasks and score their answers. Salesforce is building that same idea for AI.
The patent describes a system that sends real customer relationship management (CRM) tasks, things like answering a support question or looking up an account, to AI agents running on different AI models. Each response gets graded against a set of benchmarks, and the system produces an accuracy score.
The goal is to give companies a consistent, repeatable way to compare which AI model or which AI-powered agent actually performs best on the kinds of work that comes up inside sales and customer service software. Instead of just trusting marketing claims, businesses could run the test themselves and see the numbers.
How the benchmark scores each AI agent response
The patent covers a system that automates the process of evaluating AI agents (software programs powered by large language models, or LLMs, that can take actions on a user's behalf) specifically for tasks that come up inside CRM platforms like Salesforce's own products.
Here's the basic flow:
- The system takes a data set tied to a specific agent task (for example, "summarize this customer complaint and suggest a next step") and bundles it with a prompt.
- That prompt gets sent to one or more LLMs, which generate a response through the agent application.
- An algorithm then grades the response against a library of pre-built benchmarks, each designed to test a specific capability relevant to CRM work.
- The system outputs a numerical score measuring how accurate or appropriate the response was.
The key word in the claim is agent use case. This isn't just testing whether an AI can answer a trivia question. The benchmarks are specifically designed around the kinds of multi-step, real-world tasks that an AI agent inside a CRM would actually need to handle, like routing a ticket, drafting a follow-up email, or pulling data from a customer record.
What this means for businesses buying AI-powered CRM tools
Right now, most companies evaluating AI tools for their customer service or sales teams have to rely on vendor benchmarks, which are, predictably, flattering to the vendor. A standardized, internal scoring system built into a CRM platform would give Salesforce's enterprise customers a way to run their own tests on real tasks from their own data, before committing to a particular AI model or agent setup.
This also positions Salesforce as the neutral party setting the grading rules inside its own ecosystem, which is a significant form of platform control. If your company wants to swap in a different AI model inside Salesforce, Salesforce's benchmarking system would be the one telling you whether the new model is actually better for your specific workload.
This patent is less about a flashy new AI capability and more about who controls the measuring stick. Salesforce building a proprietary benchmark layer into its CRM platform is a quiet but meaningful power move: it gives Salesforce the authority to define what 'good' looks like for AI agents running inside its software, which is exactly the kind of infrastructure advantage that compounds over time.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
14 drawing sheets from US 2026/0228103 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Editorial commentary on a publicly published patent application. Not legal advice.