Microsoft Patents an AI System for Mapping and Modernizing Old Codebases
Every large company has a basement full of software that nobody fully understands anymore. Microsoft has filed a patent for an AI system that reads that old code, builds a map of what it all does, and answers questions about it so engineers can finally bring it up to date.
How Microsoft's AI reads and organizes aging code
Ever tried to find one email in a ten-year archive? Now imagine that archive is a million lines of software code written by people who left the company years ago, with no index and no instruction manual. That's what 'legacy code' looks like, and modernizing it is one of the most expensive headaches in the industry.
Microsoft's patent describes a system that feeds that old code into a large language model (the same kind of AI that powers tools like Copilot) to do two things automatically: build a structured map of what every part of the software does, and then create a searchable index so engineers can ask plain-English questions about the code and get useful answers.
The goal is to take software that feels like an archaeological dig and turn it into something an AI can help actively explain and update, without requiring a team of developers to read every line by hand first.
generating, using a codebase as input, a hierarchical representation of the codebase, wherein the hierarchical representation of the codebase includes one or more summarizations of functional units within the codebase and information describing relationships between the functional units within the codebase …
Translation: The AI reads old software to build a structured outline summarizing how different parts connect.
How the system builds its searchable code map
The patent describes a two-step process. First, the system takes an existing codebase as input and generates what it calls a hierarchical representation: essentially a structured tree that breaks the software down into its functional units (think: individual functions, modules, classes) and captures how they relate to each other. Each node in that tree gets an AI-generated summary describing what that chunk of code actually does in plain terms.
Second, the system converts those summaries and relationships into vector representations (numerical encodings that capture meaning, not just keywords) and stores them in a searchable index. When an engineer queries the system, the LLM searches that index to find the most relevant parts of the codebase and generates a response grounded in what the code actually contains.
The practical effect is that the LLM doesn't have to read millions of lines of raw code every time someone asks a question. Instead, it queries a pre-built map, which makes responses faster and more accurate.
- Parses the codebase and identifies discrete functional units
- Summarizes each unit and records dependencies between units
- Converts summaries into vector embeddings for semantic search
- Routes queries through the index so the LLM answers from structured context, not raw text
… generating a searchable index that includes vector representations for nodes in the hierarchical representation of the codebase, wherein the LLM generates a response to a query based on content contained in the searchable index.
Translation: It builds a smart database so the language model can accurately answer questions about the code.
What this means for companies stuck on legacy software
For large organizations, legacy code is a real financial problem. Banks, insurers, government agencies, and manufacturers often run on software written in the 1980s or 1990s that no one alive fully understands. Modernizing it typically requires months of expensive manual analysis before a single line can safely be changed. A system that automates the 'read and map' phase could cut that cost significantly.
For everyday users, the downstream effect is software that breaks less often and can be updated faster. Microsoft's interest in AI-assisted developer tooling is already visible in products like GitHub Copilot, and a code-mapping system like this could plug directly into that ecosystem as a preprocessing step before any AI-assisted rewriting begins.
Microsoft's 41st filing in the Enterprise AI patents we've tracked since May adds to a pattern that includes one reading your screen and one using your camera.
The pipeline described here runs entirely on software that already exists: tools for reading code, tools for finding meaning in text, and tools for answering questions. No new chip, no new model, no research breakthrough required before someone could start building this.
The real question is whether the map this system draws of an old codebase is detailed enough to matter. Reading a 40-year-old banking program is easy compared to knowing which parts of it are holding up the ceiling and which parts nobody has touched since 1987. The patent describes the structure of that map but cannot guarantee its accuracy on the messiest real-world systems.
If Microsoft ships something like this inside its existing developer tools, enterprise customers with aging software would have a practical reason to pay attention immediately.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
12 drawing sheets from US 2026/0299940 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in