IBM Patents a Two-Layer System for Tracing Where AI-Generated Text Came From
When an AI model spits out a paragraph, who owns it? IBM has filed a patent for a system that automatically checks generated text against a library of source documents to find out where the writing might have come from.
How IBM's text attribution system works in plain terms
Today, if you want to know whether a piece of text was copied or adapted from something else, you largely have to do that comparison by hand or with blunt plagiarism tools that miss subtle matches. IBM wants to change that with a more layered, automated approach.
The system works in two stages. First, it converts both the text you're checking and a library of reference documents into a kind of numerical fingerprint. Those fingerprints let the system quickly rule out documents that are clearly unrelated. Then, for the ones that survive that first cut, it does a closer, word-level comparison to score how similar they really are.
The practical goal is attribution: figuring out whether a piece of content, especially something produced by an AI, draws heavily enough from a source document that the source deserves credit. That's a question courts, publishers, and AI companies are all wrestling with right now.
an indexing component that generates and stores a first text representation and a first vector representation of a reference dataset …
Translation: It builds a searchable database of reference texts using both words and mathematical vectors.
Inside IBM's vector-filter-then-text-score pipeline
The patent describes a system built around three software components that work in sequence.
First, an indexing component takes a library of reference documents and generates two representations of each one: a text representation (a structured form of the actual words) and a vector representation (a list of numbers that captures the document's meaning in a way a computer can compare quickly). Think of the vector as a point on a map: documents with similar meaning land near each other.
Second, a query processing component does the same thing to the target content you want to check. It builds both a text representation and a vector representation of that content.
Third, a matching component runs a two-stage comparison:
- Vector-based filtering: it uses the numerical fingerprints to quickly narrow the reference library down to documents that are meaningfully close in topic and phrasing.
- Text-based similarity scoring: it then does a finer-grained comparison on that smaller pool, using the actual words to produce a similarity score and flag potential attribution.
The combination matters because pure vector search is fast but imprecise, while pure text matching is accurate but slow at scale. Chaining them addresses both problems.
… vector-based filtering and applies text-based similarity scoring to identify potential attribution against the designated target content …
Translation: It combines quick math filters with deep text comparisons to spot the original source of the writing.
What this means for AI copyright and content tracing
AI-generated content is flooding courts and newsrooms with the same basic question: did this model borrow too heavily from someone's work? Existing tools are mostly built for catching student plagiarism, not for tracing the subtler ways a large language model might reproduce or paraphrase a source. A system like the one IBM describes here could slot into editorial workflows, AI auditing pipelines, or legal discovery processes.
For you as a reader or creator, the downstream effect is about accountability. If this kind of attribution technology becomes standard, AI companies could be required to show their work, and content creators would have a cleaner path to claiming their material was used without permission.
IBM's 58th filing we've tracked since May in the AI safety controls race adds to a pattern that includes one on live collaboration security rules and one on access control during bug fixes.
Claim 1 is written broadly. It covers any system that (a) stores vector and text representations of a reference dataset, (b) generates matching representations for a query, and (c) uses vector filtering followed by text-based scoring to find attribution. There are no constraints on the size of the dataset, the type of AI model involved, or what counts as a sufficient similarity score. That breadth means the claim, if granted, could touch a wide range of content-verification tools, not just IBM's own.
The two-stage architecture is real engineering, and the problem it targets is genuinely pressing. But the underlying ideas, running vector search for rough filtering and then applying text similarity for precision, are well-established in information retrieval. The question examiners will ask is whether combining them specifically for attribution adds something new enough to patent.
IBM's interest in AI governance and content tracing makes this filing a coherent move. Whether the claim survives prior art review at this level of generality is a separate matter entirely.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
8 drawing sheets from US 2026/0300349 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →
Be the first to weigh in