New Google Patents · Filed Jun 23, 2025 · Published Jul 30, 2026 · verified — real USPTO data

Google Patents a System Where Two AIs Invent Fake Conversations to Train Themselves

What if you could train an AI to hold better conversations by having it practice with another AI first? That's exactly what Google is patenting here: a system where two language models talk to each other, and the transcript of that fake conversation becomes new training material.

Google Patent: AI Models Training Each Other via Fake Conversations — figure from US 2026/0220172 A1
Figure from the official USPTO publication.
See all 9 drawings from this filing ↓
Publication number US 2026/0220172 A1
Applicant Google LLC
Filing date Jun 23, 2025
Publication date Jul 30, 2026
Inventors Harsh Lara, Luke Beck Friedman, Sameer Ahuja, Manoj Kumar Tiwari
CPC classification 704/9
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (May 13, 2026)
Parent application is a National Stage Entry of PCTUS2022053882 (filed 2022-12-22)
Document 20 claims

What Google's AI-talks-to-AI training trick actually does

Training an AI to talk like a human normally requires huge amounts of real human conversation data, which is expensive and slow to collect. Google's patent describes a shortcut: instead of waiting for humans to chat, just have two AI models talk to each other and record everything they say.

One AI plays the role of a person asking questions or starting topics. The other AI plays the role of someone responding. The system captures their back-and-forth as a structured conversation, then feeds that transcript back in to train one or both of the AIs involved.

Think of it like a flight simulator for chatbots. Pilots don't need to crash a real plane to learn what to do, and with this system, an AI doesn't need millions of real human conversations to get better at having them. The synthetic conversations stand in for the real thing.

How the two-model conversation loop generates training data

The patent describes a pipeline with two generative language models (AI systems that produce human-readable text) assigned distinct roles in a simulated conversation.

  • A first model receives a conversation topic and generates an opening message, acting like a user initiating a chat.
  • A second model responds to that message, acting like a chatbot or assistant.
  • The exchange continues in ordered turns, producing a structured back-and-forth transcript called a synthetic conversational dataset.

That dataset is then used to train, pre-train, or fine-tune either the same models that generated it or entirely different ones. Fine-tuning means taking a model that already has general language ability and nudging it to behave better in a specific context, like customer support or a recommendation system.

The claim is broad enough to cover any conversational application: chatbots, recommendation engines, or interactive assistants. The key technical move is that the data generator and the data consumer can be the same model, creating a self-improving feedback loop without requiring human-written examples at every step.

Why self-generated training data is a big deal for chatbots

Real conversation data is one of the hardest things to collect at scale. Human transcripts raise privacy concerns, licensing costs, and quality inconsistencies. If Google can reliably substitute synthetic conversations for real ones, the cost of training and updating conversational AI drops substantially, and the speed at which new chatbot products can be tuned for specific use cases increases.

For you as a user, this has a direct downstream effect on products like Google Assistant, Bard (now Gemini), and any Google-powered chatbot embedded in third-party apps. A system that can rapidly generate its own training data can also update faster when conversation styles change, slang shifts, or new topic domains need to be added, without waiting months for humans to generate and label new examples.

Editorial take

This is a genuinely consequential patent in AI infrastructure, not a flashy consumer feature. The ability to generate your own training data in a closed loop is something every major AI lab is racing to figure out, and Google staking a patent claim on a specific two-model conversation architecture matters. That said, the concept of using AI to generate AI training data is already widely practiced, so how defensible this specific filing turns out to be is a real open question.

The drawings

9 drawing sheets from US 2026/0220172 A1 · click any drawing to enlarge

Patent filing page

Which company should we read for you?

We track 17 companies here. Pro is the same weekly breakdown for any company you choose, delivered privately. Type a name and we'll scope it and send you a quote.

Get one Big Tech patent every Sunday

Plain English, intelligent commentary, no hype. Free.

Source. Full patent text and figures from the official USPTO publication PDF.

Editorial commentary on a publicly published patent application. Not legal advice.