Amazon · Filed Mar 27, 2025 · Published Oct 1, 2026

Amazon Patents a Way to Stop AI Chatbots From Copying Your Documents Word for Word

AI systems built to answer questions from your own files have a quiet problem: they sometimes spit back your documents almost verbatim. Amazon has filed a patent for a way to prevent that, and the approach is more clever than a simple word limit.

A user device interacts with computing resources and large language models, which access sharded data sources to generate results. Drawing from patent filing US 2026/0300623 A1.
A user device interacts with computing resources and large language models, which access sharded data sources to generate results.
See all 11 drawings from this filing ↓
Publication number US 2026/0300623 A1
Applicant Amazon Technologies, Inc.
Filing date Mar 27, 2025
Publication date Oct 1, 2026
Inventors Aws Albarghouthi, Tancrede Lepoint
US classification 704/9
Examiner DUGDA, MULUGETA TUJI (Art Unit 2653)
Status when we published Waiting for an examiner (Jul 11, 2025)
Document 20 claims

What Amazon's document-sharding copy protection actually does

Ever tried to get a quick answer from a pile of documents, only to get a wall of text copied straight from one of them? That is the problem Amazon is trying to solve here.

When a business uses an AI assistant to answer questions about its own files, the AI can sometimes lean too hard on one specific document and repeat it back almost word for word. That is a copyright and confidentiality headache. Amazon's approach is to split the source documents into separate chunks, called shards, and run the AI's question against each chunk independently. The AI gives two sets of candidate word choices, and a combined score is calculated to pick the final word. Because no single chunk is ever the sole input, the AI cannot just parrot any one piece of text.

The idea is that if you spread the source material across multiple independent lookups, no individual passage has enough pull to drag the output into a verbatim copy. The result is a response that draws on all your documents without reproducing any one of them.

From the filing · CLAIM 1
… shard the data source into at least a first shard and a second shard; receive a request for generative text to be provided in part by a large language model (LLM) implemented with retrieval-augmented generation (RAG) and having access to at least the first shard and the second shard; …

Translation: The system chops a document into pieces before an AI model uses them to write text.

How splitting data sources keeps the AI from over-copying

This patent describes a system built around a technique called Retrieval-Augmented Generation (RAG), which is what happens when an AI language model is given access to a library of documents at question time rather than just relying on its training. The risk with RAG is that the model can get so anchored to one retrieved passage that it essentially transcribes it.

Amazon's fix works at the level of individual word choices. When the AI needs to pick its next word (a token, in technical terms), the system runs two parallel queries: one against the first shard of source material and one against the second. Each query returns a probability score for every candidate word. The system then combines those probabilities and picks the word with the highest blended score.

  • Source documents marked as sensitive are split into at least two shards.
  • Both shards are queried separately for each word the AI generates.
  • The probability scores from each shard are merged to select the final word.
  • The finished response is assembled word by word using those merged scores.

Because the two shards contain different slices of the source text, a phrase that scores highly against one shard is diluted when averaged against the other. That dampening effect makes it statistically harder for a long run of copied words to survive into the final output.

From the filing · THE ABSTRACT
These systems and processes reduce a risk that the LLM is too closely tied to any specific data and thus replicates an extended amount of that data.

Translation: This method stops the chatbot from accidentally copying source documents word for word.

What this means for businesses using AI on private documents

Businesses increasingly use AI to answer questions from internal documents, legal filings, and proprietary research. The legal exposure from an AI that copies those documents too closely is real, whether the concern is copyright, trade-secret leakage, or data-privacy regulation. A mechanism baked into the generation process itself, rather than a filter bolted on afterward, is a more reliable place to catch the problem.

For anyone building or buying an enterprise AI service, this kind of protection could become a standard expectation. Amazon's long bet on enterprise AI infrastructure shows up in filings like this one, where the goal is making AI outputs safe enough for regulated industries. If Amazon ships this in a product like Amazon Bedrock, it could let companies use RAG-based AI on sensitive documents with less legal anxiety.

Amazon's seventh filing we've tracked since July in the AI guardrails race adds to a pattern that includes one catching its own errors and one watching for suspicious inputs.

Editorial take

Claim 1 covers any system that splits a document source into pieces, runs a separate AI query against each piece, and then combines the resulting word-choice odds to decide what the AI actually says. The claim does not specify how big the pieces must be, what formula combines the odds, or which AI model is involved. That is a very wide fence.

In practice, that breadth means any competing system that shards source documents and merges word probabilities to avoid verbatim copying would fall inside this claim, regardless of how differently the details are implemented.

For everyday users, the effect is invisible and useful: you ask an AI to summarize your contracts and you get a real answer instead of a copied paragraph. But the underlying method making that happen, if this claim is granted, would belong to Amazon across a broad range of implementations.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

11 drawing sheets from US 2026/0300623 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.