Microsoft Patents a System That Sends AI Requests to Models Based on Answer Length
Asking an AI for a short summary and asking it for a full report cost the computer very different amounts of work. Microsoft's filing sorts requests by that difference before any chip starts running.
What Microsoft's answer-length dispatcher does
Someone asks an AI tool to summarize a long document. A moment later, someone else asks it to write a full report from a single sentence. The first request is heavy on reading but light on writing. The second is the opposite, and writing is the expensive part for the computer chips doing the work.
Microsoft's patent application describes a dispatcher that sits in front of many AI models. Before passing your request along, it estimates how long the answer will be, then picks a model that has room to handle it. The goal is that a quick chat reply doesn't get stuck behind a giant writing job.
For you, that could mean fewer slow replies when an AI feature gets busy. The filing also describes separate waiting lines for requests where you are staring at the screen, like a chatbot, and background jobs like summaries.
… estimating a plurality of estimated lengths of generated outputs for executing the AI request by the dedicated GPUs, estimating the plurality of estimated lengths of the generated outputs being based at least in part on a hardware characteristic of the dedicated GPUs; …
Translation: The system guesses how long the AI's answer will be using hardware traits.
How the dispatcher picks a model and chip
The patent covers an AI communication interface, the front door that apps use to send requests to AI models. When a request arrives, the system first works out which type of model can handle it, then narrows to the group of models of that type. In this setup, each model has its own graphics processors (GPUs, the chips that run AI) assigned ahead of time, before any request shows up.
The central step is an estimate of how long the generated answer will be. That estimate is based at least in part on a hardware characteristic of the chips involved. Length can be counted in tokens (small chunks of text, roughly word pieces). The request is then routed to a chosen model based on those estimates, and the answer goes back to whoever asked. The description gives examples: a summary around 50 words, a chatbot reply around 10, a document written from a subject 500 to 1,000 words or more.
Other claims add:
- An estimate of the load: computing resources, time, or memory the answer will need.
- Priority levels, with separate queues for synchronous requests (someone is waiting, like a chatbot) and asynchronous ones (background work, like summaries).
- Pulling extra data out of the prompt, such as context, and sending it along.
- Saving each prompt and response to a data store.
The description also mentions optional extras: screening for prompt injection (prompts built to make a model misbehave), rate limits, and failover if a model goes down.
… identifies whether the generative AI request is an asynchronous or a synchronous request and identifies a likely length of generation requested by the generative AI request.
Translation: The system checks if the request needs an instant reply and how long the answer will be.
Why busy AI servers need a traffic cop
AI features inside apps, such as asking for a summary of today's emails, all land on a limited supply of expensive chips. Spreading requests well across those chips can mean quicker replies for you and lower cost for the company running them. The filing says this approach lowers the computing capacity needed to serve the calls.
The design also reflects how Microsoft describes its AI setup: many apps and customers share one front door, with each model type running on its own chips. Routing by expected answer length gives that shared setup a way to keep short jobs moving while long ones take their time. You would never see it directly. You would just notice whether an AI tool stays responsive at busy hours.
Microsoft's 19th filing we've tracked since July in our AI chip wars watchlist follows one shifting storage by task and one compressing AI math tables.
Running AI costs real money, and the filing says why: writing an answer takes far more computing than reading a question. When each model has its own set of chips and requests range from a ten-word chat reply to a thousand-word document, one pool can get swamped while another sits idle. That problem costs a company serving millions of requests money every day.
The fix fits the size of the problem. Guessing the answer length before sending the job is a cheap step that can avoid a lot of wasted chip time, and it is software sitting in front of chips the company already has.
Expect it to stay out of sight. If it works, your AI tool simply feels a little less sluggish at peak hours. This one also continues an application filed in 2023, so the idea has been in the works for a while.
Get our take in your Top Stories
Liked this breakdown? Add Patentlyze as a preferred source on Google, and our plain-English take shows up more often in your Top Stories the next time Microsoft patent news breaks.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
15 drawing sheets from US 2026/0310946 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in