Salesforce · Filed Oct 23, 2025 · Published Oct 1, 2026 · verified — real USPTO data

Salesforce's New Patent Grades AI Answers on Both Meaning and Exact Wording

Getting an AI to answer questions correctly is one thing. Getting it to answer them completely, in natural language, without missing key facts, is far harder. Salesforce has filed a patent on a scoring method that judges AI answers at two levels simultaneously, then feeds that grade back into training.

A person asks an AI agent a question about a network crash, and the AI agent provides an answer. Drawing from patent filing US 2026/0300631 A1.
A person asks an AI agent a question about a network crash, and the AI agent provides an answer.
See all 10 drawings from this filing ↓
Publication number US 2026/0300631 A1
Applicant Salesforce, Inc.
Filing date Oct 23, 2025
Publication date Oct 1, 2026
Inventors Shrikant Kendre, Juan Carlos Niebles Duque, Austin Xu, Shafiq Rayhan Joty, Honglu Zhou, Michael S. Ryoo
CPC classification 704/9
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Nov 18, 2025)
Parent application Claims priority from a provisional application 63779512 (filed 2025-03-28)
Document 20 claims

How Salesforce grades AI answers to make them better

Every time a customer service chatbot gives you a wrong or half-baked answer, something failed in the way that AI was trained to know whether its responses were any good. Grading an AI's answers is surprisingly tricky: a response can sound correct without actually containing the right information, or it can include all the right facts but express them in a way that a simple comparison tool marks as wrong.

Salesforce's patent describes a system that grades AI answers on two levels at once. First, it checks whether the overall meaning of the answer matches a reference answer. Second, it checks whether the specific words and facts match. Each check produces a score, and the two scores are blended into a single grade that gets fed back to the AI so it can improve.

To make that first check work, the system first generates a "synthetic" reference answer, one that's worded the way a chatbot naturally talks, rather than a bare bullet-point fact. That way, the meaning-check doesn't unfairly penalize an AI for phrasing things conversationally. The end result is an AI that gets graded fairly and, over time, trained to be more accurate.

From the filing · CLAIM 1
… generating a first score indicative of a sentence-level semantic similarity between the synthetic answer and the generated response, based on a first similarity comparison between the embedding of the synthetic answer and the first embedding of the generated response; …

Translation: The system calculates how closely the AI answer matches the overall meaning of the target response.

How the two scores combine to train the language model

The system introduces a three-part evaluation loop for training large language models on question-answering tasks.

First, a small "synthesizer" model takes a ground-truth answer (think: the correct answer from a reference database) and rewrites it in the longer, more conversational style that a chatbot would naturally produce. This synthetic answer bridges the gap between a terse factual answer and the kind of fluent response a user would actually expect.

Second, an embedding model (software that converts text into numerical vectors representing meaning) encodes four pieces of text: the original ground-truth answer, the synthetic answer, and the AI's generated response (twice, using different encodings). From these, the system calculates two scores:

  • Sentence-level semantic score: how closely the overall meaning of the AI's response matches the synthetic answer, measured by comparing their vector representations.
  • Keyword-level score: how well the AI's response matches the ground-truth answer at the word and phrase level, using both exact matching and meaning-based keyword comparison.

Finally, the two scores are combined into a single reward signal. In reinforcement learning (a training technique where the model is rewarded for good outputs and penalized for poor ones, similar to how you'd train a dog with treats), this combined score tells the model how well it did and guides it to generate better answers in the next round.

From the filing · THE ABSTRACT
Embodiments described herein provide a QA framework that combines sentence-level semantic analysis with keyword-level semantic matching and/or exact keyword matching as feedback to iteratively update the generation and training process.

Translation: This method checks both the general meaning and specific words used to grade and improve the AI.

What better AI grading means for enterprise chatbots

For companies deploying AI assistants to answer customer questions, the quality of the answers depends almost entirely on how well the AI was trained to know when it got something right. Current scoring methods often fail at one extreme or the other: they either miss that an answer has the wrong meaning, or they penalize perfectly good answers that use different wording. A two-layer scoring approach could close both gaps at once.

Salesforce's run of LLM-training filings points to a company building the infrastructure to host and refine AI models for its enterprise clients in-house. If this scoring method works as described, it could make AI assistants in CRM and customer-service tools noticeably more reliable, without requiring a much larger or more expensive underlying model.

Salesforce files its 17th application we've tracked since May in the AI guardrails race, adding to earlier work like its agent permission system and its two-graph accuracy fix.

Editorial take

Claim 1 is broad enough to matter. It covers the entire pipeline: synthesize a verbose answer, generate two different embeddings of the model's response, compute scores at two granularities, blend them into a reward signal, and iterate. Any system training a language model on Q&A using both a sentence-level semantic score and a keyword-level score simultaneously would land squarely inside this claim.

That breadth has real consequences. Most standard evaluation pipelines use either a semantic similarity metric or an exact-match metric, not both in a single reward signal. If this patent is granted as written, it could cover a meaningful slice of how people already train Q&A models, which makes it worth watching from a competitive standpoint even if the underlying idea sounds incremental.

The genuinely clever piece is the synthetic answer generation step. By pre-translating a terse ground-truth answer into chatbot-style prose before scoring, the system avoids penalizing an AI for doing exactly what it was designed to do, which is talk like a person. That's a practical fix to a real training problem, and it keeps the two-score architecture from being just a dressed-up version of something that already exists.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

10 drawing sheets from US 2026/0300631 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.