Nvidia vs. Intel Patents on Robot Grasping, and where the race stands
This tracker collects Nvidia and Intel patent filings on teaching robots to grasp objects, navigate around obstacles, correct sensor drift, and learn from their own failures. Together the filings show two chipmakers racing to turn lab demonstrations into robots that behave reliably outside the lab.
based on all tracked filings in this watchlist · refreshes every week
These filings are all fighting over the same core problem: how do you teach a robot to see, grab, and move through the world without breaking things or getting stuck?
Nvidia and Samsung are doing the heaviest lifting here, with Nvidia focused on training robots inside simulated worlds and Samsung focused on the physical hardware that touches and rolls through the real one.
What’s new in Teaching robots to grasp and move
a dated entry each week this watchlist moves · older entries stay archived
Sep 17, 2026 1 filing joined
Sony filed a patent for robot hands that use light to sense touch more precisely. The focus this week is on giving robot fingers a finer ability to feel what they are holding.
Sony and Nvidia both filed this week, with Nvidia leading on four filings. The new work spans teaching robots to read text, feel objects, and test their own software.
This week's three new filings split across two ideas: teaching robots to understand spoken or typed object commands more flexibly, and helping robot hands feel and correct their own grip. Sony filed twice, covering both smoother movement around obstacles and better pressure-aware gripping.
Aug 20, 2026 3 filings joined
This week's filings lean heavily on Nvidia, which added two patents covering a single AI brain that can run any robot and a system that reads and labels the parts of 3D objects. Sony also joined with a patent for a two-layer artificial skin that helps robots feel touch more accurately.
Who’s filing patents in Teaching robots to grasp and move
counts from tracked filings · focus read from each company’s own filings
The battlegrounds inside Teaching robots to grasp and move
the fights inside the fight · each with its three newest filings · new filings join every week
Video As a Teaching Tool 6 filings
Nvidia 6
Several filings describe using video, including simulated and AI-generated video, to teach robots how to move and act. Nvidia is the main company pushing this approach, with Sony also filing on avoiding collisions in moving robot arms.
These filings focus on how a robot hand figures out what it is holding and adjusts each finger accordingly. Samsung and Nvidia are both filing here, covering everything from thumb design to grip pressure to finger-by-finger control.
These filings describe systems where a robot or AI keeps training itself in a circle, learning from failures and improving without stopping. Nvidia leads this group, with filings covering self-correcting loops, teacher-student setups, and round-the-clock training cycles.
These filings cover robots that notice when something has gone wrong and change their plan on the fly, without a person stepping in. Samsung and Intel are both active here, alongside Nvidia.
These filings describe robots that build and update pictures of the space around them so they can move through it safely. Samsung is the main filer, with patents covering room mapping, choosing the best viewing spot, and tracking where workers are.
These filings describe ways to tell a robot what to do using spoken words or by copying human body movements directly. Intel and Nvidia are both filing here, covering voice commands, motion capture, and breaking instructions into steps.
Earlier filings in this tracker show robots struggling to calibrate grip strength; Sony's optical approach trades strain gauges for light-based sensing, offering a path to finer pressure discrimination in robot fingers.
Robots that parse natural language commands could work in unstructured spaces without hand-coded coordinates, reducing setup time and expanding where they're useful in real warehouses.
Robots that fuse distance and contact data before committing to a grasp reduce the guesswork baked into vision-only approaches. This filing suggests sensor redundancy as the path to more reliable object handling in unstructured environments.
Controlling force along different axes independently lets robots navigate tight spaces without explicit obstacle instructions, the arm pushes hard in goal directions while yielding passively to constraints.
Testing robot software is slow, inconsistent, and often done by hand. Nvidia is filing a patent for a system that automates the whole process, running robots through tasks and scoring their performance the same way every time.
Robots could learn new tasks from synthetic training footage generated from a single real video, cutting the time needed to gather the hundreds of examples they typically require.
Robots could adjust their grip mid-motion if cameras track object movement across frames instead of just spotting static targets. This builds the continuous 3D awareness needed to handle shifts in position that static snapshots can't catch.
Robots could follow novel object commands without retraining on every variation, letting them generalize from prior grasping experience to new situations.
Sensor drift during insertion: pressure feedback across fingertips lets the robot detect misalignment and self-correct mid-motion, without waiting for mechanical failure to signal the problem.
Robots that navigate tight spaces would gain real-time responsiveness by running slimmer obstacle-detection models instead of full neural networks, cutting both training time and power draw during deployment.
Stacking soft and firm material layers lets robots distinguish touch pressure and location simultaneously, filling a gap in current grip-control systems that confuse gentle contact with forceful squeezing.
Teaching a computer to look at a 3D model of a chair and immediately know which bits are legs, which is the seat, and which is the back turns out to be surprisingly hard. Nvidia's latest patent describes an AI approach that does exactly that, breaking any 3D shape into a tree of labeled pieces.
A unified model trained on multiple body morphologies lets one AI handle humanoids, quadrupeds, and wheeled platforms without retraining from scratch, consolidating what previously required separate control systems for each robot type.
Selecting the best-performing robot AI from dozens of training iterations currently requires manual testing. Nvidia's system automates this grading by running candidate models through standardized tests in simulation and real-world conditions.
Grip adjustment during handling solves a separate problem from initial grasping: robots need feedback loops that let them shift finger pressure mid-lift when an object rotates or proves heavier than expected.
The tracker so far has shown Nvidia and Intel working on sensor drift and learning from failure. This filing adds memory of recent actions, so robots can correct course mid-task rather than repeat the same mistakes across separate attempts.
Where the watchlist has focused on finger coordination, this filing zeroes in on thumb geometry, specifically the multiple rotation axes needed to replicate how a human thumb crosses opposing fingers.
The timeline so far has focused on learning from failure; this filing shifts to real-time course correction, where the robot adjusts mid-movement based on what its camera sees rather than relying on preset coordinates.
The watchlist so far has focused on learning from failure and correcting sensor errors. This filing adds a foundation layer: getting the physics simulation itself accurate enough that training in simulation actually transfers to real robots.
Synthetic image generation lets robots learn object recognition without collecting thousands of real-world photos across different angles and lighting. This reduces the time and cost of assembling training datasets for grasping tasks.
The watchlist so far has focused on teaching robots to grasp objects and learn from failure. Sony's patent adds a real-time layer: keeping jointed arms from colliding with obstacles during movement itself, not just planning the route beforehand.
Distributing simulation across multiple data centers lets the system collect training data from parallel experiments rather than sequential runs, compressing the months of trial-and-error into weeks.
A feedback loop between two AI models lets robots refine their own training data by learning from real-world attempts, then using that data to improve the next round of learning without human labeling.
The watchlist so far has shown robots learning from failure and correcting their own errors. This filing extends that by keeping the learning cycle running continuously while the robot operates, rather than stopping to retrain between tasks.
Continuous motion blending during execution lets robots shift between actions mid-movement rather than completing one scripted gesture before starting the next, reducing the jarring transitions that plague pre-programmed routines.
Generating training video in simulation rather than capturing it from real robots sidesteps the cost and time of collecting actual footage, letting the system iterate through edge cases at scale.
Tracking how object relationships shift in real time lets robots update their next move without replanning from scratch, solving the navigation problem when environments change mid-task.
Robots can now detect when their sensors drift mid-task and recalibrate in real time, letting them keep working through unexpected physical bumps without restarting.
Generating physics-accurate video from text descriptions lets robots learn collision avoidance and object dynamics in simulation before touching real warehouses, compressing months of trial-and-error into synthetic training data.
Within the gripper-control problem, this filing shows how to run reachability and collision-checking together rather than as separate steps, cutting the decision time a robot arm needs before it moves.
A world foundation model generates predicted video of the target task, letting robots preview actions before execution, confirming the tracker's focus on learning from visual simulation rather than trial-and-error in the physical world.
Synthetic video generation skips the physical trial-and-error loop by letting an AI create realistic robot demonstrations from workspace images and task descriptions, then uses those videos as training data.
Where the watchlist has focused on learning from real failures, this filing shows robots could instead practice in physics-accurate simulations, collapsing trial-and-error into virtual space before touching actual objects.
Sensor gloves paired with headset tracking split the workload between hand-detail capture and full-body positioning, letting a human operator demonstrate complex manipulation tasks to a robot in real time rather than programming each movement.
Within the grasp-and-move watchlist, this filing shows how robots can learn placement tasks from minimal human input, moving past the hundreds-of-example bottleneck by extracting spatial logic from a single demonstration video.
A correction layer that monitors prediction errors and updates the physics model between movements, keeping the robot's control calculations aligned with real-world conditions as they drift.
Previous grasp failures often stem from misreading object geometry or category. This filing merges shape and identity recognition into one step, reducing the mismatch errors that cause awkward approaches.
A robot that removes obstacles instead of stopping at them cuts down the trial-and-error cycles needed to navigate real homes. The mechanical gripper solves the navigation problem by making clutter moveable rather than trying to route around it.
Translating natural language commands into robot motion by learning from paired video and voice narration sidesteps the need to manually code movement parameters for each task.
Robots executing fixed plans crash when sensor data degrades mid-task. Intel's filing shows how to reweight a robot's priorities on the fly, trading off task completion speed against uncertainty about what the robot can actually sense in real time.
A depth camera lets the robot scan multiple height layers at once and select whichever one contains the clearest landmarks for navigation, solving the problem of getting stuck when a single camera angle sees only clutter.
A robot that spreads its fingers before grasping could handle objects ranging from thin cards to wide bottles without reprogramming between tasks, reducing the need for multiple specialized grippers.
The patents so far show robots learning to manipulate physical objects. Samsung's filing extends the problem into the digital area: a robot must read a screen, understand what it shows, and operate touchscreens or remote interfaces as a human would.
Powered joints that enforce correct motion sequences during exercise represent a shift from passive support to active movement correction, addressing how robots learn limb positioning through physical guidance rather than visual observation alone.
A robot that corrects its own mistakes without human intervention moves the grasp problem from initial training to real-world durability, where small execution errors compound into mission failures.
The watchlist so far has focused on object manipulation; Samsung shifts the lens to planning failures when a robot's mental map proves incomplete, forcing real-time replanning rather than task abandonment.
When a robot arm gets physically stuck mid-task, it normally needs human rescue. Samsung's filing shows how onboard AI can detect these deadlock states and compute escape paths without outside intervention.
Earlier filings focused on grasping and object manipulation; this one zooms out to the locomotion layer itself, solving how wheeled robots stay on curved paths without drift by individually controlling wheel speeds based on predictive trajectory data.
A robot that stops smoothly instead of lurching to a halt could reliably set down objects without spilling or dropping them, a basic requirement for any real warehouse work.
The watchlist so far treats vision as fixed; Samsung's approach makes it mobile, letting robots reposition their own cameras to see past obstacles rather than work around blind spots.
Verification loops stand out as the missing piece: Samsung's patent adds machine vision that compares floor conditions before and after cleaning, letting robots detect when stains persist and decide whether to retry a spot.
A central monitoring system tracks worker movement in real time and feeds dynamic route adjustments directly to the robot fleet, replacing the pre-planned paths that create bottlenecks when humans cluster in one area.
A robot that reads human gaze could pick objects from cluttered spaces without needing explicit verbal commands, reducing the real-time communication overhead that slows down human-robot collaboration in actual homes and warehouses.
Simulating touch feedback without perfect object positions pushes training into territory where robots must learn from incomplete sensory data, mirroring real-world grasping constraints.
A two-stage simulation approach separates learning object geometry from learning real-world motor control, letting the robot master spatial reasoning before encountering sensory noise and delays.
The robot grasping problem requires training data that doesn't overwhelm the learning system from the start. Nvidia's approach sequences tasks by difficulty, letting the AI build competence incrementally rather than failing on hard cases too early.
Robots could infer object weight from human posture during handoffs, eliminating the need to sense weight directly at the moment of grasping and reducing grip failures from misjudged force.
Automated detection of human interventions during robot operations creates a continuous feedback loop for retraining without manual annotation, converting real-world failures into training data at scale.
The watchlist so far has focused on training robots to physically handle objects. This filing pivots to the perceptual problem: giving robots spatial memory so they can navigate without constant repositioning.
Sensor drift compounds during long warehouse runs, forcing robots to periodically recalibrate. Intel's dual-sensor fusion weights each signal's reliability to maintain continuous position accuracy without stopping.
Separating grasp point prediction from grip planning into two distinct neural networks lets each stage specialize rather than forcing one model to solve both problems simultaneously.
Diffusion models, normally used for image generation, retrain here as simulators that generate candidate grasp sequences from scratch, letting the system learn which approaches work before robots try them physically.
A robot trained in simulation can now adjust its own behavior when real-world physics don't match predictions, eliminating the costly trial-and-error phase that usually follows deployment.
Questions readers ask
What problem are Nvidia and Intel trying to solve with these robot patents?
They're mostly working on the same core problem: getting robots to reliably grasp objects and move through real environments without constant human correction. The filings show sim-to-real training loops, sensor fusion for drift, and systems that let robots learn from failed attempts. These are research filings, not confirmed products, so they show direction rather than finished technology.
Does a robot patent mean the robot is actually being built?
No. A patent filing describes an idea a company wants to protect, not a shipping product. Nvidia and Intel file broadly across grasping, navigation, and perception, and only some of these ideas will end up in real robots. The filings are a useful signal of where engineering attention is going, not a roadmap.
Why do so many of these patents focus on grasping instead of walking or other robot skills?
Grasping keeps showing up because picking up an unfamiliar object reliably is still unsolved in robotics. Nvidia has filed multiple grasping approaches, including diffusion-based and teacher-student methods, which suggests the company sees this as worth attacking from several angles at once rather than settling on one method.
What does Samsung's involvement add to this watchlist?
Samsung's filings are more mechanical than Nvidia's or Intel's, covering a two-stage braking system for rolling robots and a robot that repositions its own sensor to see around obstacles. That shows the same reliability concerns, staying in control and seeing clearly, showing up in hardware design and not just AI training methods.
Want this weekly breakdown for a company we don't cover?
Patentlyze Pro →
The weekly email: the best of Big Tech's filings, in plain English. Free.