Sony · Filed Mar 5, 2025 · Published Sep 10, 2026 · verified — real USPTO data

Sony's New Patent Covers Software That Rewrites Its Own Priorities on the Fly

Most AI systems lock in their priorities at training time and can't change course when things go wrong. Sony's new patent describes an agent that spots trouble ahead and shifts its own internal priorities to avoid it, all without being retrained.

Sony Patent: AI Agents That Rebalance Their Own Goals Mid-Task — figure from US 2026/0268206 A1
Figure from the official USPTO publication.
See all 4 drawings from this filing ↓
Publication number US 2026/0268206 A1
Applicant SONY GROUP CORPORATION
Filing date Mar 5, 2025
Publication date Sep 10, 2026
Inventors Alisa Devlic, Alfredo Reichlin, Kaushik Subramanian, Peter Wurman
CPC classification 706/12
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (Apr 1, 2025)
Document 20 claims

How Sony's self-adjusting AI agent changes its own goals

Imagine you're playing a video game against an AI opponent that can see it's about to make a fatal mistake. Instead of blindly charging ahead, it pauses, reconsiders what matters most right now, and adjusts its strategy on the spot. That's roughly what Sony's patent describes, but for AI systems in general, not just games.

Most AI agents are trained with a fixed set of goals and priorities. If those priorities turn out to be wrong for the current situation, you have to go back, retrain the agent, and try again. Sony's approach lets a single trained policy cover a wide range of priority settings. The agent can then shift those settings in real time when it predicts something is about to go wrong.

The practical result is an AI that can handle a wider variety of situations without needing a full do-over. You get more flexible behavior from one trained model, rather than needing a separate model for every possible scenario.

From the filing · CLAIM 1
… adaptively adjusting, during operation of the artificial intelligent agent, the weight for one or more selected parameters of the parameterized reward functions.

Translation: The software changes its own operational goals while it is currently running.

How the agent rebalances reward weights during operation

The patent builds on an idea called a Universal Value Function Approximator (UVFA), which is a way of training an AI agent so that its goals are treated as inputs to the model rather than as fixed constants baked in at training time. Think of it like a dial instead of a switch: rather than the agent being set permanently to "aggressive" or "cautious," you can turn the dial to any point in between.

Sony's system extends that idea to compositional reward functions, meaning the agent's total reward is broken into several components (say, speed, safety, and energy use), each with its own adjustable weight. During training, the agent is exposed to many different combinations of those weights, sampled from a continuous range. This teaches it to perform well across the entire spectrum, not just at a handful of preset points.

At runtime, the agent can:

  • Monitor incoming signals to predict a future negative event (a collision, a failure, a penalty)
  • Automatically shift the weights assigned to relevant reward components
  • Continue operating under the new priority mix without stopping or retraining

The key engineering claim is that a single policy handles all of this. You don't need a library of specialized agents for every scenario; one flexible agent covers the full range.

From the filing · THE ABSTRACT
… adaptively adjust the components' weights, in runtime, in efforts to avoid the predicted future negative event.

Translation: The system tweaks its priorities on the fly to steer clear of trouble it sees coming.

What adaptive AI priorities mean for games and robotics

For anyone playing games against or alongside Sony's AI systems, this could mean opponents or teammates that feel less mechanical and more adaptive. An AI that can recognize when its current approach is failing and shift gears is a lot harder to exploit with a single repeated tactic.

Beyond games, Sony's ongoing investment in AI agent research points toward robotics and automation, where a robot that can rebalance its priorities (say, favoring safety over speed when a person walks nearby) without requiring a complete software update would be genuinely useful. The patent doesn't guarantee any specific product, but the underlying idea of runtime priority adjustment matters in any setting where conditions change faster than a developer can retrain a model.

That makes this Sony's seventh filing in the AI assistant and agent space we've tracked since May, joining earlier work on cloning players as AI bots and mimicking human writing styles.

Editorial take

The real change for a player or user is that the AI can stop making the same frustrating mistake on repeat. Instead of waiting until something breaks down, the system anticipates the problem and shifts its priorities before the failure happens.

What that looks like in practice is an AI opponent or teammate that recalibrates mid-session, without anyone flipping a switch. A character that was too aggressive starts playing it safer, or one that ignored defense starts covering ground it previously abandoned.

How noticeable that actually feels depends entirely on how well the system reads the warning signs early enough to act on them. If that prediction piece works, the experience feels alive. If it does not, the behavior shifts will seem random rather than intelligent.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

4 drawing sheets from US 2026/0268206 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.