Microsoft's New Patent Cuts the Cost of Teaching AI to Connect Information
When AI systems search large document collections, the biggest hidden cost isn't the search itself, it's building the map of concepts the search depends on. Microsoft has filed a patent for a method that trims that upfront cost without sacrificing the quality of the answers.
What Microsoft's adaptive knowledge graph actually does
Every time an AI assistant answers a question about a large pile of documents, something expensive happens in the background: the system has to read all those documents and build a kind of map showing which ideas appear near each other. That map is what lets it give you a coherent answer instead of random snippets.
The problem is that building a high-quality map is slow and computationally heavy. Microsoft's patent describes a two-step shortcut: start by chopping the documents into larger text blocks to get a rough draft of the concept map quickly, then zero in on smaller blocks only where the map needs more detail. Instead of doing the most expensive work everywhere, you do it only where it counts.
The result is a concept map that's good enough to be useful but built at a fraction of the usual cost. That matters most for companies running AI tools over large internal document libraries, where every query has a real computing price tag.
chunking source documents into overlapping text chunks having a first size; extracting concepts in the text chunks; creating a co-occurrence graph of the concepts; and, refining the co-occurrence graph by changing to a second chunk size.
Translation: Breaking documents into overlapping pieces and building a concept map, then adjusting the piece size to improve it.
How the chunk-size shift refines the concept graph
The patent covers a technique inside a family of AI retrieval systems called GraphRAG (Graph-based Retrieval-Augmented Generation). GraphRAG works by turning a document collection into a graph, a web of concepts and the relationships between them, so an AI can answer questions by traversing that web rather than scanning raw text.
The specific method works in four steps:
- Chunking: source documents are split into overlapping text segments (so ideas that span a boundary aren't lost) at a first, larger chunk size.
- Concept extraction: the system identifies meaningful terms and entities inside each chunk.
- Co-occurrence graph creation: it records which concepts appear near each other, building a graph where edges represent "these ideas show up together."
- Refinement: the graph is then improved by switching to a second chunk size, allowing the system to recalibrate where the initial pass was too coarse or too fine.
The key insight is that chunk size drives both cost and accuracy. Larger chunks are cheaper to process but miss fine-grained relationships; smaller chunks catch nuance but multiply the computation. By starting coarse and refining selectively, the patent proposes a way to hit a better cost-quality balance than running a single fixed chunk size from start to finish.
This document relates to providing meaningful information relating to a dataset.
Translation: The technology helps make sense of large amounts of stored data.
What this means for running AI search on real data
For any organization running AI-powered search over its own documents, legal files, research archives, internal wikis, the cost of building and maintaining the underlying knowledge graph is a real operational expense. A method that cuts that cost while preserving answer quality could make GraphRAG-style systems practical at scales where they're currently too expensive to run continuously.
For everyday users, the effect would be invisible but felt: AI assistants that can answer detailed questions about large document sets faster and at lower operating cost, which matters when those costs determine whether a product feature gets built at all. Microsoft's bet on GraphRAG is visible across multiple filings, and this one targets the part of the pipeline that most affects whether the technology stays affordable as document collections grow.
Microsoft's 72nd filing in our Language AI coverage since May adds to a run that includes one on image meaning search and one grouping cloud alerts.
Building a working knowledge map across thousands of documents is expensive enough that the cost alone decides whether many organizations can attempt it at all. That is a real barrier, not a technical footnote.
Microsoft's answer is to run the map-building process in two rounds, adjusting how documents get sliced up between them rather than locking in one approach from the start. The logic is sound: doing less work upfront and refining later should cut the bill meaningfully.
The harder question is whether those two rounds stay well-calibrated when the documents are messy and varied, as they almost always are outside a controlled setting. A cost-saving method that only works cleanly on clean material solves a smaller problem than the one it is advertising.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
32 drawing sheets from US 2026/0300370 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in