Qualcomm Patents an AI That Grades Its Own Moves to Control Robots
Most robots follow a script. Qualcomm's new patent describes an AI that checks its own work after each step and rewrites the rest of the plan if something went wrong.
How Qualcomm's self-grading robot AI actually works
Ever watched someone try to assemble furniture without the instructions? They put a piece in, realize it's wrong, and figure out a new approach. Qualcomm's patent describes an AI that does something similar for robots: try a step, look at the result, score it, and decide what to do next.
The idea is that you give the robot a goal, say "sort these packages into two bins." The AI picks a first action, carries it out, then compares where things ended up to where they were supposed to be. That comparison produces a score. A bad score means the robot tries a different approach; a good score means it keeps going.
This matters because most automated systems are brittle. They follow a fixed sequence, and if anything goes sideways, they stop or fail silently. A system that self-corrects mid-task is much closer to how a human worker actually handles surprises on the job.
… calculating a reward metric based on a comparison of a post-execution state of an environment in which the autonomous device is operating and a target state of the environment in which the autonomous device is operating …
Translation: The system checks how close the robot got to its goal by comparing the room before and after it moved.
How the reward metric drives the replanning loop
The patent describes a closed-loop control system built around a generative AI model (a large model capable of producing text, images, or action sequences, not just classifying inputs). The model receives a multimodal input, meaning it can take in text instructions, images, sensor readings, or a mix, and uses that to identify a goal.
Once the goal is set, the system runs a three-phase loop:
- Plan: The model identifies a first sub-task to execute toward the goal.
- Act: The autonomous device carries out that sub-task in the real (or simulated) environment.
- Evaluate: The model calculates a reward metric (a numerical score measuring how close the post-action world state is to the intended target state). Based on that score, it picks one or more follow-on tasks.
The loop repeats until the goal is reached or the model determines it cannot proceed. The reward metric is the key mechanism: it replaces a static checklist with a continuous judgment call after every action.
The patent also covers training this kind of model, not just running it. That means Qualcomm is describing both how to build the system and how to deploy it, which is a broader claim than most robotics-control filings.
Certain aspects of the present disclosure provide techniques and apparatus for controlling an autonomous device, such as a robot, using a machine learning model and for training such a model.
Translation: This patent covers a method for using and training artificial intelligence to operate robots.
What this means for real-world robot reliability
For you as an end user, the practical difference is reliability. A robot that can only follow a fixed script fails the moment the environment doesn't match expectations, a box is in the wrong spot, a tool slips, a step takes longer than expected. A system that scores its own performance and adjusts has a fighting chance of recovering without a human stepping in.
Qualcomm's track record in edge-AI and on-device compute patents suggests this is aimed at running on hardware that goes inside devices, not in a cloud data center. If this kind of self-correcting control loop can run locally on a chip, it becomes practical for drones, warehouse robots, or mobile devices that can't always phone home for instructions.
Qualcomm's eighth filing we've tracked in our AI assistant & agent filings coverage since June adds to earlier work on keeping only key video frames and stopping AI video memory overload.
The cost is time. After every action, the robot pauses to grade its own work before deciding what comes next. In a warehouse that pause is invisible, but in a burning building it could be the difference between useful and dangerous.
The deeper risk is that the grading itself can be wrong. If the robot's definition of "success" is even slightly off, it will keep confidently doing the wrong thing, with no outside check to catch it.
That said, building correction directly into every step is cleaner than bolting a recovery system onto a robot never designed to recover. Whether the tradeoff holds depends entirely on how fast and how honest that self-grading turns out to be once it leaves the lab.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
11 drawing sheets from US 2026/0295819 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in