Salesforce's New Patent Grades AI Answers on Both Meaning and Exact Wording
Getting an AI to answer questions correctly is one thing. Getting it to answer them completely, in natural language, without missing key facts, is far harder. Salesforce has filed a patent on a scoring method that judges AI answers at two levels simultaneously, then feeds that grade back into training.
How Salesforce grades AI answers to make them better
Every time a customer service chatbot gives you a wrong or half-baked answer, something failed in the way that AI was trained to know whether its responses were any good. Grading an AI's answers is surprisingly tricky: a response can sound correct without actually containing the right information, or it can include all the right facts but express them in a way that a simple comparison tool marks as wrong.
Salesforce's patent describes a system that grades AI answers on two levels at once. First, it checks whether the overall meaning of the answer matches a reference answer. Second, it checks whether the specific words and facts match. Each check produces a score, and the two scores are blended into a single grade that gets fed back to the AI so it can improve.
To make that first check work, the system first generates a "synthetic" reference answer, one that's worded the way a chatbot naturally talks, rather than a bare bullet-point fact. That way, the meaning-check doesn't unfairly penalize an AI for phrasing things conversationally. The end result is an AI that gets graded fairly and, over time, trained to be more accurate.
… generating a first score indicative of a sentence-level semantic similarity between the synthetic answer and the generated response, based on a first similarity comparison between the embedding of the synthetic answer and the first embedding of the generated response; …
Translation: The system calculates how closely the AI answer matches the overall meaning of the target response.
How the two scores combine to train the language model
The system introduces a three-part evaluation loop for training large language models on question-answering tasks.
First, a small "synthesizer" model takes a ground-truth answer (think: the correct answer from a reference database) and rewrites it in the longer, more conversational style that a chatbot would naturally produce. This synthetic answer bridges the gap between a terse factual answer and the kind of fluent response a user would actually expect.
Second, an embedding model (software that converts text into numerical vectors representing meaning) encodes four pieces of text: the original ground-truth answer, the synthetic answer, and the AI's generated response (twice, using different encodings). From these, the system calculates two scores:
- Sentence-level semantic score: how closely the overall meaning of the AI's response matches the synthetic answer, measured by comparing their vector representations.
- Keyword-level score: how well the AI's response matches the ground-truth answer at the word and phrase level, using both exact matching and meaning-based keyword comparison.
Finally, the two scores are combined into a single reward signal. In reinforcement learning (a training technique where the model is rewarded for good outputs and penalized for poor ones, similar to how you'd train a dog with treats), this combined score tells the model how well it did and guides it to generate better answers in the next round.
Embodiments described herein provide a QA framework that combines sentence-level semantic analysis with keyword-level semantic matching and/or exact keyword matching as feedback to iteratively update the generation and training process.
Translation: This method checks both the general meaning and specific words used to grade and improve the AI.
What better AI grading means for enterprise chatbots
For companies deploying AI assistants to answer customer questions, the quality of the answers depends almost entirely on how well the AI was trained to know when it got something right. Current scoring methods often fail at one extreme or the other: they either miss that an answer has the wrong meaning, or they penalize perfectly good answers that use different wording. A two-layer scoring approach could close both gaps at once.
Salesforce's run of LLM-training filings points to a company building the infrastructure to host and refine AI models for its enterprise clients in-house. If this scoring method works as described, it could make AI assistants in CRM and customer-service tools noticeably more reliable, without requiring a much larger or more expensive underlying model.
Salesforce files its 17th application we've tracked since May in the AI guardrails race, adding to earlier work like its agent permission system and its two-graph accuracy fix.
Claim 1 is broad enough to matter. It covers the entire pipeline: synthesize a verbose answer, generate two different embeddings of the model's response, compute scores at two granularities, blend them into a reward signal, and iterate. Any system training a language model on Q&A using both a sentence-level semantic score and a keyword-level score simultaneously would land squarely inside this claim.
That breadth has real consequences. Most standard evaluation pipelines use either a semantic similarity metric or an exact-match metric, not both in a single reward signal. If this patent is granted as written, it could cover a meaningful slice of how people already train Q&A models, which makes it worth watching from a competitive standpoint even if the underlying idea sounds incremental.
The genuinely clever piece is the synthetic answer generation step. By pre-translating a terse ground-truth answer into chatbot-style prose before scoring, the system avoids penalizing an AI for doing exactly what it was designed to do, which is talk like a person. That's a practical fix to a real training problem, and it keeps the two-score architecture from being just a dressed-up version of something that already exists.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
10 drawing sheets from US 2026/0300631 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in