Microsoft · Filed Jun 11, 2026 · Published Oct 1, 2026

Microsoft Patents a System That Flags Suspicious AI Prompts Before They Cause Damage

Every time someone types a prompt into a company's AI assistant, there's a chance they're trying to manipulate it. Microsoft has filed a patent for a system that learns what "normal" looks like for an AI model and its users, then fires off a security alert when something looks off.

A network connects multiple user devices and servers, with one server containing logic for profile-based anomaly detection. Drawing from patent filing US 2026/0300473 A1.
A network connects multiple user devices and servers, with one server containing logic for profile-based anomaly detection.
See all 8 drawings from this filing ↓
Publication number US 2026/0300473 A1
Applicant Microsoft Technology Licensing, LLC
Filing date Jun 11, 2026
Publication date Oct 1, 2026
Inventors Aviv SHITRIT, Roee OZ, Idan HEN, Tamer SALMAN, Alon DANOCH, Ron KELLER, Asaf HARARI
US classification 726/23
Status when we published Waiting for an examiner (Jun 28, 2026)
Parent application is a Continuation of 18939462 (filed 2024-11-06)
Document 20 claims

How Microsoft's AI watchdog spots unusual chat behavior

Today, most AI security tools focus on scanning prompts for specific forbidden words or known attack patterns. The problem is that sophisticated attacks rarely look obviously bad on the surface. Microsoft wants to change that by watching patterns of behavior over time, not just individual messages.

The idea is to build two kinds of profiles: one for the AI model itself (tracking what kinds of conversations it normally has and what kinds of answers it normally gives), and one for each user (tracking how they typically phrase questions and what sessions usually look like). When an incoming message looks dramatically different from those established patterns, the system can step in and take a security action, like blocking the prompt or flagging it for review.

Think of it as a fraud-detection system for AI chat. Your bank notices when a purchase happens in a country you've never visited. This does the same thing, but for prompts sent to a company's AI tools.

From the filing · THE ABSTRACT
Techniques are described herein that are capable of performing a security action based on anomaly detection using AI model profiles and user profiles.

Translation: The patent outlines a method for catching suspicious behavior by comparing prompts against established behavior baselines.

How the system builds profiles and measures divergence

The patent describes a system that generates and maintains four types of behavioral profiles:

  • Model-session profile: captures the typical semantic content (meaning, topic, and tone) of full conversations with the AI model across many sessions.
  • Model-response profile: tracks what kinds of answers the model normally produces.
  • User-session profiles: one per user, representing what a typical back-and-forth conversation looks like for that individual.
  • User-prompt profiles: one per user, representing how that user normally phrases their individual requests.

When a new prompt arrives, the system computes how far it deviates from one or more of these profiles. "Semantic meaning" here means the system is looking at the intent and topic of a message, not just its keywords, likely using an embedding model (a kind of AI that converts text into numerical coordinates, where similar meanings land close together in space).

If the difference between the incoming prompt and the stored profiles crosses a threshold, the system triggers a security action. That action could be blocking the request, requiring additional authentication, logging it for a human reviewer, or something else the platform operator defines.

The design is notable because it guards against two different threat vectors: someone from outside acting like an unusual user, and an AI model that starts behaving in ways it never did before (which could indicate the model itself has been tampered with or manipulated through prior prompts).

What this means for companies running AI tools at work

For anyone whose company uses an AI assistant (for customer support, internal search, code generation, or anything else), this kind of system addresses a real and growing risk. Attackers are increasingly trying to "jailbreak" or manipulate enterprise AI tools by feeding them carefully crafted prompts. Traditional content filters miss attacks that use benign-sounding language to lead the AI somewhere it shouldn't go. A pattern-based approach catches the unusual behavior even when the words look innocent.

This also matters at the model level, not just the user level. a growing pile of Microsoft AI-security filings shows the company thinking about cases where the AI's own response patterns shift unexpectedly, which is a harder problem and a less-discussed threat surface than user-side attacks.

Microsoft's 21st filing we've tracked in our AI safety guardrails work since July builds on earlier applications like one reviewing medical AI chats and one catching bias in legal analysis.

Editorial take

If you use an AI assistant at work, the real threat this addresses is one you'd never catch on your own: a bad actor nudging a conversation across many small exchanges until the AI is doing something it shouldn't. This system watches the whole conversation as it unfolds, so that slow drift gets flagged before it causes damage.

It also builds a behavioral baseline for the AI itself. If the tool starts responding in ways that fall outside its normal patterns, that's treated as a warning signal even when the person typing looks completely legitimate.

The honest catch is that these baselines require time and real usage to become reliable. A team that just adopted a new AI tool may have thin protection in the first few weeks, which is exactly when mistakes are most likely.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

8 drawing sheets from US 2026/0300473 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.