US2017211206A1PendingUtilityA1

Methods, systems, and software for identifying bio-molecules with interacting components

Assignee: CODEXIS INCPriority: Jan 31, 2013Filed: Apr 4, 2017Published: Jul 27, 2017
Est. expiryJan 31, 2033(~6.5 yrs left)· nominal 20-yr term from priority
Inventors:Gregory A. Cope
G16B 5/00G16B 20/00G16C 20/60G16B 15/00G16B 35/00G16C 10/00C12N 15/1058G06F 19/701G06F 19/12G06F 19/16G06F 19/18C40B 50/02G16C 20/50G16B 10/00G16C 20/30G01N 33/50G16B 35/20G16B 35/10G16B 5/20G16B 20/50G16B 20/20
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present invention provides methods for rapidly and efficiently searching biologically-related data space. More specifically, the present invention provides methods for identifying bio-molecules with desired properties, or which are most suitable for acquiring such properties, from complex bio-molecule libraries or sets of such libraries. The present invention also provides methods for modeling sequence-activity relationships, including but not limited to stepwise addition or subtraction techniques, Bayesian regression, ensemble regression and other methods. The present invention further provides digital systems and software for performing the methods provided herein.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, implemented using a computer system comprising one or more processors and system memory, of conducting directed evolution of bio-molecules, the method comprising:
 (a) receiving, by the computer system, sequence data and activity data for a plurality of bio-molecules;   (b) fitting, by the one or more processors, a first or second base model to the sequence data and activity data, wherein
 the first or second base model relates sub-units of a sequence of a bio-molecule to an activity of the bio-molecule, 
 the first base model includes one or more linear terms but no interaction term, 
 the second base model includes one or more linear terms and a defined pool of interaction terms, and 
 each interaction term represents an interaction between two or more interacting sub-units; 
   (c) obtaining, by the one or more processors, a plurality of new models, wherein each new model is obtained by adding to the first base model one different interaction term in the defined pool of interaction terms or subtracting from the second base model one different interaction term in the defined pool of interaction terms;   (d) determining, by the one or more processors, an ability of each model of the plurality of new models to predict the activity as a function of the presence or absence of the sub-units;   (e) identifying, by the one or more processors, at least one best model among the plurality of new models based on the ability of the plurality of new models to predict activity as determined in (d) and with a bias against including additional interaction terms;   (f) identifying, using the at least one best model, one or more identified bio-molecules to modify in a round of directed evolution.   
     
     
         2 . The method of  claim 1 , wherein obtaining the plurality of new models in (c) comprises using prior information of parameters of the plurality of new models to determine posterior probability distributions of the parameters of the plurality of new models. 
     
     
         3 . The method of  claim 2 , wherein the fitting a first or second base model and/or obtaining the plurality of new models comprises using Gibbs sampling to fit a model to the sequence and activity data. 
     
     
         4 . The method of  claim 1 , wherein the at least one best model comprises two or more best models, each of which includes different interaction terms. 
     
     
         5 . The method of  claim 1 , further comprising preparing an ensemble model based on the two or more best models, wherein
 the ensemble model includes interaction terms from the two or more best models, and   the interaction terms are weighted by the ability of the two or more best models to predict activity as determined in (d).   
     
     
         6 . The method of  claim 1 , further comprising:
 (i) setting the at least one best model as an updated model and repeating (c) using the updated model in place of the first or second base model; and   (ii) repeating (d) and (e).   
     
     
         7 . The method of  claim 6 , further comprising repeating (i) and (ii) one or more times. 
     
     
         8 . The method of  claim 1 , wherein the ability of the plurality of new models to predict the activity in (d) is measured by Akaike Information Criterion or Bayesian Information Criterion. 
     
     
         9 . The method of  claim 1 , wherein the sequence is a whole genome, whole chromosome, chromosome segment, a collection of gene sequences for interacting genes, gene, or protein. 
     
     
         10 . The method of any of  claim 1 , wherein the sub-units are chromosomes, chromosome segments, haplotypes, genes, nucleotides, codons, mutations, amino acids, or residues. 
     
     
         11 . The method of  claim 1 , wherein the plurality of bio-molecules comprises a protein variant library. 
     
     
         12 . The method of  claim 1 , wherein (f) comprises identifying the one or more identified bio-molecules that are predicted by the at least one best model to have a desired level of activity. 
     
     
         13 . The method of  claim 12 , further comprising performing saturation mutagenesis on the one or more identified bio-molecules that are predicted by the at least one best model to have the desired level of activity. 
     
     
         14 . The method of  claim 1 , wherein the each of the one or more identified bio-molecules comprises sub-units associated with one or more coefficients of the at least one best model, and wherein the one or more coefficients meet one or more criteria. 
     
     
         15 . The method of  claim 1 , wherein (f) comprises:
 evaluating coefficients of the one or more linear terms and/or coefficients of the one or more interaction terms of the at least one best model to identify one or more defined amino acids at defined sequence positions that contribute to the activity;   selecting one or more mutations of the defined amino acids; and   identifying one or more oligonucleotides encoding the one or more mutations.   
     
     
         16 . The method of  claim 15 , wherein the evaluating coefficients comprises identifying one or more coefficients that are determined to be larger than other coefficients. 
     
     
         17 . The method of  claim 1 , wherein the one or more identified bio-molecules comprise one or more nucleic acid molecules, and wherein the method further comprises synthesizing the one or more nucleic acid molecules using a nucleic acid synthesizer. 
     
     
         18 . The method of  claim 1 , further comprising fragmenting and recombining the one or more identified bio-molecules. 
     
     
         19 . A computer system comprising one or more processors and system memory, the one or more processors being configured to:
 (a) receive sequence data and activity data for a plurality of bio-molecules;   (b) fit a first or second base model to the sequence data and activity data, wherein
 the first or second base model relates sub-units of a sequence of a bio-molecule to an activity of the bio-molecule, 
 the first base model includes one or more linear terms but no interaction term, 
 the second base model includes one or more linear terms and a defined pool of interaction terms, and 
 each interaction term represents an interaction between two or more interacting sub-units; 
   (c) obtain a plurality of new models, wherein each new model is obtained by adding to the first base model one different interaction term in the defined pool of interaction terms or subtracting from the second base model one different interaction term in the defined pool of interaction terms;   (d) determine an ability of each model of the plurality of new models to predict the activity as a function of the presence or absence of the sub-units;   (e) identify at least one best model among the plurality of new models based on the ability of the plurality of new models to predict activity as determined in (d) and with a bias against including additional interaction terms; and   (f) identify, using the at least one best model, one or more identified bio-molecules to modify in a round of directed evolution.   
     
     
         20 . A computer program product comprising a non-transitory machine readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to implement a method for conducting directed evolution of bio-molecules, said program code comprising:
 (a) code for receiving sequence data and activity data for a plurality of bio-molecules;   (b) code for fitting a first or second base model to the sequence data and activity data, wherein
 the first or second base model relates sub-units of a sequence of a bio-molecule to an activity of the bio-molecule, 
 the first base model includes one or more linear terms but no interaction term, 
 the second base model includes one or more linear terms and a defined pool of interaction terms, and 
 each interaction term represents an interaction between two or more interacting sub-units; 
   (c) code for obtaining a plurality of new models, wherein each new model is obtained by adding to the first base model one different interaction term in the defined pool of interaction terms or subtracting from the second base model one different interaction term in the defined pool of interaction terms;   (d) code for determining an ability of each model of the plurality of new models to predict the activity as a function of the presence or absence of the sub-units;   (e) code for identifying at least one best model among the plurality of new models based on the ability of the plurality of new models to predict activity as determined in (d) and with a bias against including additional interaction terms; and   (f) code for identify, using the at least one best model, one or more identified bio-molecules to modify in a round of directed evolution.

Join the waitlist — get patent alerts

Track US2017211206A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.