Microsoft · Filed Mar 26, 2026 · Published Jul 30, 2026 · verified — real USPTO data

Microsoft Patents AI System That Cites Sources and Rejects Off-Topic Questions

One of the most persistent problems with AI assistants is that they confidently answer questions they have no business answering. Microsoft is patenting a training method that teaches AI to say 'that's outside my area' instead of guessing.

Microsoft Patent: Domain-Specific AI That Knows What It Doesn't Know — figure from US 2026/0220141 A1
Figure from the official USPTO publication.
See all 10 drawings from this filing ↓
Publication number US 2026/0220141 A1
Applicant Microsoft Technology Licensing, LLC
Filing date Mar 26, 2026
Publication date Jul 30, 2026
Inventors Atabak ASHFAQ, Haiyuan CAO, Yu Hu
CPC classification 707/722
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Apr 24, 2026)
Parent application is a Continuation of 18988766 (filed 2024-12-19)
Document 20 claims

What Microsoft's domain-limited AI actually does

Imagine you hire a contractor who specializes in plumbing. You'd want them to fix your pipes and tell you upfront when you ask about electrical work that it's out of their lane. That's essentially what Microsoft is trying to build for AI.

This patent describes a way to train smaller, specialized AI models that are tuned to a specific industry or company's knowledge. When you ask the AI something it knows about, it answers and cites where that answer came from. When you ask something outside its designated area, it declines rather than making something up.

The system uses two types of training data working together: one set that teaches the AI what to answer and how to cite sources, and another that teaches it to recognize the same question even when it's phrased differently. The goal is a reliable AI assistant that a business can actually trust with customer-facing or internal tasks.

How the two-dataset training process works

The patent describes a training pipeline for what Microsoft calls a domain-integrated contextual response model, a smaller AI tuned to a specific company or industry's knowledge base.

The process starts with a synthetic data generator that creates two types of practice questions: in-domain questions (things the AI should know) and out-of-domain questions (things outside its scope). The system then runs those questions through a more powerful AI model, like a large language model (LLM), to generate high-quality example answers. This is called skill distillation (learning how to answer from a bigger model) combined with knowledge distillation (learning what to answer from the company's own data).

Two training datasets are built from this process:

  • First dataset: teaches the model to answer in-domain questions with cited sources and to refuse out-of-domain questions entirely
  • Second dataset: uses paraphrased versions of the same questions, so the model isn't thrown off by different wording

The result is a smaller model that can be deployed by a specific business, handles Retrieval-Augmented Generation (RAG) tasks (meaning it pulls answers from a real document store rather than pure memorization), and stays reliably within its lane.

What this means for enterprise AI tools like Copilot

For businesses deploying AI, the biggest risk isn't what the AI doesn't know. It's what the AI doesn't know it doesn't know and answers anyway. A customer service bot that confidently gives wrong legal or medical information creates real liability. This patent addresses that directly by building refusal behavior into the training itself, not just bolting it on as a filter afterward.

Microsoft's Copilot products are already pushing into enterprise workflows across finance, healthcare, and legal sectors. A system like this would let those customers deploy an AI that behaves like a narrow specialist rather than a generalist that wanders into dangerous territory. That's a meaningful selling point for regulated industries.

Editorial take

This is exactly the kind of unglamorous infrastructure patent that determines whether enterprise AI is actually usable or just a liability. Teaching an AI to reliably say 'I don't know' is harder than it sounds, and Microsoft patenting a systematic training approach for it suggests they're serious about making Copilot trustworthy in high-stakes industry deployments. Worth watching.

The drawings

10 drawing sheets from US 2026/0220141 A1 · click any drawing to enlarge

Patent filing page

Which company should we read for you?

We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.

Get one Big Tech patent every Sunday

Plain English, intelligent commentary, no hype. Free.

Source. Full patent text and figures from the official USPTO publication PDF.

Editorial commentary on a publicly published patent application. Not legal advice.