IBM · Filed Mar 31, 2025 · Published Oct 1, 2026 · verified — real USPTO data

IBM Patents a Way to Teach AI the Hidden Structure Inside Drug Molecules

Drug molecules can be written out as text strings, but most AI models treat them like random letter sequences. IBM has filed a patent for a training technique that forces the model to actually learn which groups of characters form meaningful chemical structures.

A detailed molecular structure and its corresponding SMILES representation, used to teach AI about drug molecules. Drawing from patent filing US 2026/0301884 A1.
A detailed molecular structure and its corresponding SMILES representation, used to teach AI about drug molecules.
See all 9 drawings from this filing ↓
Publication number US 2026/0301884 A1
Applicant International Business Machines Corporation
Filing date Mar 31, 2025
Publication date Oct 1, 2026
Inventors Thanh Lam Hoang, Raúl Fernández Díaz, Mykhaylo Zayats, Vanessa Lopez Garcia
CPC classification 702/19
Grant likelihood Medium
Examiner CENTRAL, DOCKET (Art Unit OPAP)
Status Docketed New Case - Ready for Examination (May 2, 2025)
Document 25 claims

How IBM's AI reads a molecule like a sentence

A chemist stares at a string of letters like CC(=O)Oc1ccccc1C(=O)O, that's aspirin, written in a shorthand called SMILES. To you or me, it looks like a password. To a trained chemist, specific clusters of those characters describe recognizable chemical shapes that determine how the molecule behaves.

The problem is that AI models often miss those clusters too. They process the string character by character, without understanding that certain characters far apart in the text are actually part of the same physical structure in the molecule. IBM's patent describes a way to fix that during training: the model is shown a molecule string, then forced to guess which masked-out sections belong together as a meaningful substructure.

The goal is AI that can reliably predict drug-relevant properties, like whether a compound is toxic, how it binds to a target protein, or how the body processes it. That kind of prediction is one of the expensive, slow steps in pharmaceutical research that IBM's long bet on AI-assisted drug discovery is trying to speed up.

From the filing · CLAIM 1
… masking, using a machine learning model, all tokens in the sequence string representation corresponding to a chemical substructure within the molecular graph of the chemical molecule using a plurality of discontinuous substrings in the sequence string representation …

Translation: The system hides specific parts of the molecular text so the AI has to figure out what is missing.

How the masking and hash-vector training loop works

The patent describes a training pipeline built around a concept borrowed from language models: masked prediction. In standard language-model training, you hide a word and ask the model to guess it. IBM's approach hides entire chemical substructures instead.

Here is how it works step by step:

  • A molecule is first converted into a SMILES or SELFIES string, two standardized text notations that encode atomic bonds as characters.
  • The system identifies a chemically meaningful substructure (think: a benzene ring, an amine group) within the molecule's graph representation.
  • All characters in the string that correspond to that substructure are masked out simultaneously, even if those characters are not sitting next to each other in the text.
  • Instead of asking the model to guess each masked character separately, the model must predict a single hash vector, a compact numeric fingerprint representing the entire set of masked characters at once.

The key insight is that chemical substructures are defined by connectivity in the molecule's three-dimensional graph, not by position in the text string. A ring structure might scatter its characters across the string with other atoms in between. By masking all of them together and training the model to recognize them as a unit, the system teaches the transformer to understand higher-order relationships (meaning: interactions between multiple parts of the molecule, not just adjacent atoms).

The result is a model that builds richer internal representations of molecules, which should translate into more accurate predictions of drug-related properties.

From the filing · THE ABSTRACT
… transformers learning high-order interaction between important substructures of molecules, and thus enhancing the ability of those transformers to learn and predict drug-related properties …

Translation: The AI learns how different molecular pieces interact to better predict how drugs will behave.

What this means for AI-driven drug discovery

Drug discovery is extraordinarily expensive, and a large fraction of that cost comes from testing compounds that turn out to be toxic, inactive, or impossible for the body to absorb. AI models that can predict those properties from a molecule's structure, before any lab work begins, could cut years and billions of dollars from early-stage research. The catch is that those models are only as useful as their ability to genuinely understand molecular structure, not just pattern-match on text.

For you as a reader, the practical stakes are straightforward: better molecular AI means faster filtering of bad drug candidates, which eventually means cheaper and faster paths to treatments. This filing is about the training methodology, not a finished product, but the underlying problem it addresses is one of the most consequential open questions in pharmaceutical AI today.

IBM's 36th filing we've tracked in Language AI since May adds to a run that includes one on rewriting prompts and one on cleaner code output.

Editorial take

The problem here is real and expensive. Pharmaceutical companies spend enormous sums running wet-lab experiments on molecules that computational screening should have flagged as non-starters. The gap between what AI models claim to predict and what actually holds up in the lab is partly a representation problem: the models do not truly understand molecular structure.

IBM's approach is technically credible. Masking chemically coherent substructures rather than random tokens is a principled fix to a real limitation, and using a hash vector to represent a set of disconnected string positions is an elegant way to avoid creating a trivially easy prediction task. The core idea has analogs in recent graph-neural-network and protein-language-model research.

That said, the patent is a training method for a model component. It is a research-infrastructure filing, not a product announcement. Whether this specific technique produces meaningful accuracy gains over existing molecular transformers is an empirical question the patent naturally does not answer. Interesting work, but the proof will be in the benchmarks, not the filing.

There are more where this came from

We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.

The drawings

9 drawing sheets from US 2026/0301884 A1 · click any drawing to enlarge

Patent filing page

Source. Full patent text and figures from the official USPTO publication PDF.
Reader comments

Be the first to weigh in

Start the discussion

Real name or a handle, either is fine. Comments are read by a person before they appear, so allow a little time. Keep it about the filing.