Microsoft Patents an AI Security Agent That Has a Built-In Critic to Catch Its Own Errors
Most AI security tools act first and get reviewed by a human later. Microsoft's new patent describes an AI agent that audits its own decisions in real time, using an internal critic to catch mistakes before they cause harm.
How Microsoft's self-auditing AI security agent works
Today, automated security tools can take actions on a network, like blocking a login or flagging a file, without any real-time check on whether those actions are correct. Microsoft wants to change that by building a referee directly into the AI itself.
The system described in this patent uses three cooperating AI roles. An assistant figures out what to do when a security event happens. A user proxy carries out the assistant's instructions and gathers results. Then a critic reviews those results and flags any errors before anything is finalized. If the critic finds a problem, the feedback loops back to the assistant so it can try again.
Think of it like a surgeon, a nurse who runs the procedure, and a second surgeon who checks the work before the patient is closed up. The goal is fewer unchecked automated actions in your security system, and more confidence that the AI's conclusions are actually right.
… the agent comprising an assistant, a user proxy, and a critic, wherein the user proxy is configured to receive information related to an event and implement an agentic flow to determine one or more tools to process the event …
Translation: The AI system splits its work between three internal roles that handle tasks, run tools, and check for mistakes.
Inside the assistant-proxy-critic conversation loop
The patent describes an agentic flow (a chain of automated AI decisions and actions) structured as a three-part team inside a single AI agent.
- Assistant: receives a security event, recommends which software tool to use to investigate or respond to it, and later writes up a report on what the tool returned.
- User proxy: the central coordinator. It receives the security event, asks the assistant for guidance, runs the recommended tool, and sends the result report to the critic for review.
- Critic: reads the tool's result report, checks for errors (misclassifications, incomplete actions, or bad outputs), and sends any problems back through the user proxy to the assistant.
This plays out in two sequential conversations. First, the assistant and user proxy go back and forth to pick and run a tool. Second, the user proxy and critic review what happened. If the critic detects errors, the loop restarts. The framework is designed to integrate with threat mappings (structured databases of known attack patterns, like MITRE ATT&CK) so the AI's decisions are grounded in established security knowledge.
The whole architecture is meant to reduce the chance that an autonomous AI agent takes a wrong action in a live security environment without anyone noticing.
The critic reviews the tool response report for errors and returns any detected errors to the user proxy. The user proxy relays the detected errors to the assistant.
Translation: A built-in reviewer checks the work and sends any mistakes back to the helper to fix.
What self-checking AI means for security automation
Security automation is already happening at scale. Companies use AI-driven tools to respond to threats in milliseconds, far faster than any human team could. But when those tools make mistakes, the errors can be hard to catch and costly to fix. A miscategorized alert can mean a real attack gets ignored, or a legitimate user gets locked out.
By folding a checking mechanism directly into the AI agent, Microsoft is tackling the accountability gap that comes with autonomous systems. For you as an end user, this kind of design could mean fewer false alarms and fewer missed threats in enterprise software products. Microsoft has been filing around agentic AI safety since at least early 2025, and this patent fits a pattern of treating AI autonomy as something that needs guardrails built in, not bolted on later.
Microsoft's 30th filing we've tracked since May in our AI models working in teams builds on earlier applications including one that catches its own errors and one that rewrites its own prompts.
Using AI to check AI only works if the reviewer is meaningfully different from the system it is reviewing. If both share the same blind spots, the reviewer will confidently approve the same mistakes, every time.
The bigger cost is speed. Running two full rounds of automated conversation before acting on a security threat takes real time, and some attacks are over before that second review completes. That tradeoff, more careful for slower, is probably acceptable for routine security work but becomes a liability when the window to respond is measured in seconds.
The core idea of building a review step directly into the process, rather than hoping a human catches problems later, is a defensible design. Whether it earns its overhead depends entirely on how distinct and capable that reviewer role is in practice, and this document does not fully answer that question.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
9 drawing sheets from US 2026/0303627 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in