US2026051370A1PendingUtilityA1

Generative design of small molecules

Assignee: GENENTECH INCPriority: Mar 10, 2023Filed: Sep 5, 2025Published: Feb 19, 2026
Est. expiryMar 10, 2043(~16.6 yrs left)· nominal 20-yr term from priority
G16C 20/70G16C 20/50G06N 20/00G06N 3/08G16C 20/90G16C 20/40G06N 3/00G16C 20/30
68
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods for the generative design of small molecules include performing, until one or more conditions are satisfied, one or more iterations of a generative algorithm. Each iteration of the generative algorithm may include modifying one or more molecules from an initial population of molecules. Moreover, each iteration of the generative algorithm may include selecting, from the initial population of molecules and the one or more modified molecules, a quantity of molecules satisfying one or more fitness scores for inclusion in a subsequent population of molecules. If the one or more conditions are not satisfied, one or more additional iterations of the generative algorithm may be performed using a different initial population of molecules or the subsequent generation of molecules as a new initial population of molecules. Related systems and computer program products are also provided.

Claims

exact text as granted — not AI-modified
1 . A computer-implemented method, comprising:
 (a) obtaining, from a molecular structure database, an initial population of molecules;   (b) generating a first new population of molecules by modifying at least a first molecule in the initial population of molecules and selecting, from the initial population of molecules and at least the first modified molecule, a quantity of molecules satisfying one or more fitness scores;   (c) generating a second new population of molecules by modifying at least a second molecule in the first new population of molecules and selecting, from the first new population of molecules and at least the second modified molecule, another quantity of molecules satisfying the one or more fitness scores;   (d) in response to determining that a convergence criterion has not been satisfied, repeating step (c) at least once with the second new population of molecules as a first new population of molecules; and   (e) in response to determining that the convergence criterion has been satisfied, selecting a subset of the second new population of molecules as candidates for synthesis and testing.   
     
     
         2 . The method of  claim 1 , wherein:
 the generating the first new population of molecules comprises:   (i) calculating one or more fitness scores for at least the first modified molecule and for each molecule of the initial population of molecules; and   (ii) forming the first new population of molecules by selecting, from the initial population of molecules and at least the first modified molecule, the quantity of molecules having one or more fitness scores that satisfy the one or more fitness scores;   and   each instance of the generating the second new population of molecules comprises:   (i) calculating one or more fitness scores for at least the second modified molecule and each molecule in the first new population of molecules; and   (ii) forming the second new population of molecules by selecting, from the first new population of molecules and at least the second modified molecule, the another N quantity of molecules having one or more fitness scores that satisfy the one or more fitness scores.   
     
     
         3 . The method of  claim 2 , wherein the modifying each of the first molecule and the second molecule comprises:
 applying to the structure of the molecule one or more modifications selected from the group consisting of: substituting a non-hydrogen atom radical of the molecule with a moiety selected from a standard set of moieties; substituting a hydrogen atom of the molecule with a moiety selected from the standard set of moieties; replacing a divalent fragment of the molecule with a divalent moiety selected from the standard set of moieties, and wherein:   a radical comprises from 1-5 non-hydrogen atoms or is a ring system comprising 3-10 non-hydrogen atoms;   a moiety comprises from 1-5 non-hydrogen atoms or is a ring system comprising 3-10 non-hydrogen atoms; and   a divalent fragment comprises from 1-5 non-hydrogen atoms or is a ring system comprising 3-10 non-hydrogen atoms.   
     
     
         4 . The method of  claim 2 , wherein the one or more fitness scores comprise a diversity score. 
     
     
         5 . The method of  claim 4 , wherein the diversity score is indicative of a chemical similarity between pairs of molecules, and wherein the diversity score penalizes at least one molecule in a pair of molecules that are structurally or chemically similar to one another. 
     
     
         6 . The method of  claim 4 , wherein the diversity score is calculated from a metric selected from the group consisting of: a Tanimoto index, a cosine coefficient, a Dice metric, an Euclidean metric, a city-block metric, a Hamming index, and a Tversky index. 
     
     
         7 . The method of  claim 2 , wherein the one or more fitness scores comprises a calculated docking score for a molecule against a target binding site. 
     
     
         8 . The method of  claim 1 , wherein the convergence criterion is selected from:
 a fixed number of repeats of step (c);   an average fitness score that satisfies a threshold value, wherein the average fitness score is calculated for the second new population of molecules obtained from the last instance of step (c); and   a designated number of molecules in the second new population of molecules obtained from the last instance of step (c) has a fitness score that is less than a threshold value.   
     
     
         9 . The method of  claim 1 , wherein the obtaining the initial population of molecules comprises:
 clustering, based on a similarity metric, a corresponding set of molecules in the molecular structure database into one or more clusters of molecules; and   selecting, from each of the one or more clusters of molecules, one or more molecules having a fitness score satisfying one or more initial fitness scores.   
     
     
         10 . The method of  claim 9 , wherein the similarity metric is a Tanimoto index, a cosine coefficient, a Dice metric, an Euclidean metric, a city-block metric, a Hamming index, or a Tversky index, and wherein the similarity metric is based on a property selected from:
 2D similarity, 3D similarity, and a vector of physicochemical properties.   
     
     
         11 . The method of  claim 9 , wherein the clustering is performed by applying one or more clustering algorithms selected from: Butina, centroid, CLink, Gower, McQuitty, SLink, Unweighted Pair Group Method with Arithmetic Mean (UPMGA), Ward, and Jarvis-Patrick. 
     
     
         12 . The method of  claim 2 , wherein the one or more fitness scores for a molecule includes a calculated value of one or more of the following molecular properties: solubility, permeability, a selectivity score, an efficiency score, toxicity, and a physiologically based pharmacokinetic (PBPK) score. 
     
     
         13 . The method of  claim 1 , wherein the initial population of molecules is obtained by a method chosen from:
 randomly selecting one or more molecules from the molecular structure database;   selecting one or more molecules from the molecular structure database that have one or more physicochemical properties that meet threshold criteria; and   selecting one or more molecules that have a particular scaffold;   or a combination thereof.   
     
     
         14 . The method of  claim 1 , further comprising:
 (f) synthesizing at least one of the candidate molecules.   
     
     
         15 . The method of  claim 2 , wherein the second new population of molecules is generated by selecting, as at least a portion of the another N quantity of molecules, a fixed proportion of molecules originating from the first new population of molecules. 
     
     
         16 . The method of  claim 3 , wherein the one or more modifications for each molecule are identified by:
 determining a set of applicable modifications for the molecule from a list of available modifications, and   randomly choosing one or more modifications from the set of applicable modifications.   
     
     
         17 . A computer-implemented method, comprising:
 (a) obtaining, from a molecular structure database, a plurality of initial populations of molecules;   (b) generating a plurality of first new populations of molecules from each of the plurality of respective initial populations of molecules, wherein each of the plurality of first new populations of molecules comprises one or more molecules obtained by modifying one or more molecules in the initial population of molecules on which the first new population is based, wherein the molecules in each of the plurality of first new populations of molecules satisfy one or more fitness scores;   (c) generating a plurality of second new populations of molecules from each of the respective plurality of first new populations of molecules, wherein each of the plurality of second new populations comprise one or more molecules obtained by modifying one or more molecules in the first new population of molecules on which the second new population is based, wherein the molecules in each of the plurality of second new populations of molecules satisfy the one or more fitness scores;   (d) in response to determining that a convergence criterion has not been satisfied, repeating step (c) at least once with each of the first new populations of molecules being one of the second new populations of molecules obtained from the prior instance of step (c); and   (e) in response to determining that the convergence criterion has been satisfied, selecting a subset of molecules that comprises at least one molecule from each of the plurality of second new populations of molecules for synthesis and testing.   
     
     
         18 . The method of  claim 17 , wherein the one or more fitness scores comprise a diversity score. 
     
     
         19 . The method of  claim 18 , wherein the diversity score is based on a structural similarity between a pair of molecules that comprises one molecule selected from each of two populations of molecules, and wherein the diversity score penalizes at least one molecule in a pair of molecules that are structurally similar to one another. 
     
     
         20 . The method of  claim 18 , wherein the diversity score is calculated from a metric selected from the group consisting of: a Tanimoto index, a cosine coefficient, a Dice metric, an Euclidean metric, a city-block metric, a Hamming index, and/or a Tversky index. 
     
     
         21 . The method of  claim 18 , wherein the diversity score is calculated between every pair of molecules that can be formed from molecules in one population and molecules in another population. 
     
     
         22 . The method of  claim 18 , wherein for any pair of molecules whose diversity score does not satisfy a threshold, only one molecule from the pair will satisfy the one or more fitness scores. 
     
     
         23 . The method of  claim 17 ,
 the generating the plurality of first new population of molecules comprises performing the following for each first new population of molecules:   (i) calculating one or more fitness scores for at least the first modified molecule and for each molecule of the corresponding initial population of molecules; and   (ii) forming the first new population of molecules by selecting, from the corresponding initial population of molecules and at least the first modified molecule, an N quantity of molecules having one or more fitnesss score that satisfy the one or more fitness scores;   and   each instance of the generating the plurality of second new population of molecules comprises performing the following for each second new population of molecules:   (i) calculating one or more fitness scores for at least the second modified molecule and each molecule in the corresponding first new population of molecules; and   (ii) forming the second new population of molecules by selecting, from the corresponding first new population of molecules and at least the second modified molecule, another N quantity of molecules having one or more fitness scores that satisfy the one or more fitness scores.   
     
     
         24 . A computer-implemented method, comprising:
 (a) obtaining an initial population of molecules wherein each molecule has a molecular structure constructed from a reaction database that contains a set of reactions, wherein each reaction from the set of reactions is associated with two or more sets of reagents, and wherein each molecule in the initial population of molecules is a product of a first reagent from a first set of reagents and a second reagent from a second set of reagents, and optionally a third reagent from a third set of reagents, in accordance with a reaction to which both the first and second reagents and the optional third set of reagents are associated;   (b) generating a first new population of molecules by
 modifying at least a first molecule in the initial population of molecules by replacing the first reagent from which the first molecule is formed with another first reagent from the first set of reagents associated with the reaction from which the first molecule is formed, and optionally replacing the second reagent from which the first molecule is formed with another second reagent from the second set of reagents associated with the reaction from which the first molecule is formed, thereby forming a first modified molecule, and 
 selecting, from the initial population of molecules and at least the first modified molecule, a quantity of molecules that satisfy one or more fitness scores for inclusion in the first new population of molecules; 
   (c) generating a second new population of molecules by at least
 modifying at least a second molecule in the first new population of molecules by replacing the first reagent from which the second molecule is formed with another first reagent from the first set of reagents, and optionally replacing the second reagent from which the second molecule is formed with another second reagent from the second set of reagents thereby forming a second modified molecule, and 
 selecting, from the first new population of molecules and at least the second modified molecule, another quantity of molecules that satisfy the one or more fitness scores; 
   (d) in response to determining that a convergence criterion has not been satisfied, repeating step (b) at least once with the second new population of molecules as a new initial population of molecules; and   (e) in response to determining that the convergence criterion has been satisfied, selecting a subset of the second new population of molecules as candidates for synthesis and testing.   
     
     
         25 . A computer-implemented method, comprising:
 (a) obtaining, from a molecular structure database, an initial population of molecules, each molecule of the initial population of molecules being a product of a first reaction between a first reagent selected from a first set of reagents associated with the first reaction and a second reagent selected from a second reagent selected from a second set of reagents associated with the first reaction, and optionally a third reagent from a third set of reagents associated with the first reaction;   (b) generating a first new population of molecules by
 creating a first modified molecule by applying a modification to at least a first molecule in the initial population of molecules, wherein the modification is selected from one or more of:
 replacing the first reagent from which the first molecule is formed with another first reagent selected from the first set of reagents; 
 replacing the second reagent from which the first molecule is formed with another second reagent selected from the second set of reagents; and 
 replacing a third reagent from which the first molecule is formed with another third reagent selected from the third set of reagents; 
 
 and 
 selecting, from the initial population of molecules and at least the first modified molecule, a quantity of molecules satisfying one or more fitness scores; 
   (c) generating a second new population of molecules by
 creating a second modified molecule by applying a modification to at least a second molecule in the second new population of molecules, wherein the modification is selected from one or more of:
 replacing the first reagent from which the first molecule is formed with another first reagent selected from the first set of reagents; 
 replacing the second reagent from which the first molecule is formed with another second reagent selected from the second set of reagents; and 
 replacing a third reagent from which the first molecule is formed with another third reagent selected from the third set of reagents; 
 
 and 
 selecting, from the first new population of molecules and at least the second modified molecule, another N quantity of molecules satisfying the one or more fitness scores; 
   (d) in response to determining that a convergence criterion has not been satisfied, repeating step (c) at least once with the second new population of molecules as a new initial population of molecules; and   (e) in response to determining that the convergence criterion has been satisfied, selecting a subset of the second new population of molecules as candidates for synthesis and testing.   
     
     
         26 . The method of  claim 25 , wherein the generating a first new population of molecules further comprises:
 modifying an additional first molecule in the initial population of molecules by:
 selecting a third reagent that has a calculated similarity to the first reagent from which the additional first molecule is formed, wherein the third reagent is in an additional first set of reagents associated with an additional reaction, wherein the additional reaction is different to the reaction from which the additional first molecule is formed; 
 constructing an additional first modified molecule by replacing the first reagent from which the additional first molecule is formed, with the third reagent, and replacing the second reagent from which the additional first molecule is formed with a fourth reagent selected from an additional second set of reagents associated with the additional reaction, wherein the fourth reagent has a calculated similarity to the second reagent; 
 and 
 selecting, from the initial population of molecules, the first modified molecule, and the additional first modified molecule, a quantity of molecules satisfying one or more fitness scores. 
   
     
     
         27 . The method of  claim 26 , wherein the generating a second new population of molecules further comprises:
 modifying an additional second molecule in the initial population of molecules by:
 selecting a third reagent that has a calculated similarity to the first reagent from which the additional second molecule is formed, wherein the third reagent is in an additional first set of reagents associated with an additional reaction, wherein the additional reaction is different to the reaction from which the additional second molecule is formed; 
 constructing an additional second modified molecule by replacing the first reagent from which the additional second molecule is formed with the third reagent, and replacing the second reagent from which the additional second molecule is formed, with a fourth reagent selected from a fourth additional second set of reagents associated with the additional reaction, wherein the fourth reagent has a calculated similarity to the second reagent; 
 and 
 selecting, from the first new population of molecules, the second modified molecule, and the additional second modified molecule, a quantity of molecules satisfying one or more fitness scores. 
   
     
     
         28 . A system, comprising:
 at least one data processor; and   at least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising:   (a) obtaining, from a molecular structure database, an initial population of molecules;   (b) generating a first new population of molecules by modifying at least a first molecule in the initial population of molecules and selecting, from the initial population of molecules and at least the first modified molecule, a quantity of molecules satisfying one or more fitness scores;   (c) generating a second new population of molecules by modifying at least a second molecule in the first new population of molecules and selecting, from the first new population of molecules and at least the second modified molecule, another quantity of molecules satisfying the one or more fitness scores;   (d) in response to determining that a convergence criterion has not been satisfied, repeating step (c) at least once with the second new population of molecules as a first new population of molecules; and   (e) in response to determining that the convergence criterion has been satisfied, selecting a subset of the second new population of molecules as candidates for synthesis and testing.   
     
     
         29 . A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising:
 (a) obtaining, from a molecular structure database, an initial population of molecules;   (b) generating a first new population of molecules by modifying at least a first molecule in the initial population of molecules and selecting, from the initial population of molecules and at least the first modified molecule, a quantity of molecules satisfying one or more fitness scores;   (c) generating a second new population of molecules by modifying at least a second molecule in the first new population of molecules and selecting, from the first new population of molecules and at least the second modified molecule, another quantity of molecules satisfying the one or more fitness scores;   (d) in response to determining that a convergence criterion has not been satisfied, repeating step (c) at least once with the second new population of molecules as a first new population of molecules; and   (e) in response to determining that the convergence criterion has been satisfied, selecting a subset of the second new population of molecules as candidates for synthesis and testing.

Join the waitlist — get patent alerts

Track US2026051370A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.