US2005084907A1PendingUtilityA1
Methods, systems, and software for identifying functional biomolecules
Est. expiryMar 1, 2022(expired)· nominal 20-yr term from priority
Inventors:Richard J. Fox
G06F 30/00G16B 35/00G01N 33/6818G16B 30/00C12N 15/1058G16C 20/60G16B 20/00G16B 40/00G06N 20/00G16B 35/20G16B 20/50G16B 40/20G16B 30/10G16B 20/20
48
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The present invention generally relates to methods of rapidly and efficiently searching biologically-related data space. More specifically, the invention includes methods of identifying bio-molecules with desired properties, or which are most suitable for acquiring such properties, from complex bio-molecule libraries or sets of such libraries. The invention also provides methods of modeling sequence-activity relationships. As many of the methods are computer-implemented, the invention additionally provides digital systems and software for performing these methods.
Claims
exact text as granted — not AI-modified1 . A method for identifying amino acid residues for variation in a protein variant library in order to affect a desired activity, said method comprising:
(a) receiving data characterizing a training set of a protein variant library, wherein the data provides activity and sequence information for each protein variant in the training set; (b) from the data, developing a sequence-activity model that predicts activity as a function of amino acid residue type and corresponding position in a protein sequence, wherein the sequence-activity model includes one or more non-linear terms, each representing an interaction between two or more amino acid residues in the protein sequence; and (c) using the sequence-activity model to identify one or more amino acid residues at specific positions for variation to impact the desired activity.
2 . The method of claim 1 , wherein at least one of the non-linear terms is a cross-product term comprising a product of one variable representing the presence of one interacting residue and another variable representing the presence of another interacting residue.
3 . The method of claim 2 , wherein the sequence-activity model comprises a sum of said at least one cross-product term and one or more linear terms, each representing the presence of a variable residue in the training set.
4 . The method of claim 2 , wherein developing said sequence-activity model comprises selecting one or more cross-product terms from a group of potential cross-product terms.
5 . The method of claim 4 , wherein selecting the one or more cross-product terms comprises running a genetic algorithm to select a cross-product terms based upon the predictive ability of various models employing different cross-product terms.
6 . The method of claim 1 , wherein the protein variants in the protein variant library have systematically varied sequences.
7 . The method of claim 6 , further comprising performing DOE to identify the systematically varied sequences.
8 . The method of claim 1 , further comprising:
(d) using the sequence activity model to identify one or more amino acid residues that are to remain fixed in a new protein variant library.
9 . The method of claim 1 , wherein the protein variant library comprises naturally occurring proteins or proteins derived therefrom.
10 . The method of claim 9 , wherein the naturally occurring proteins comprise proteins that are encoded by members of a single gene family.
11 . The method of claim 1 , wherein the protein variant library comprises proteins that are obtained by using a recombination-based diversity generation mechanism.
12 . The method of claim 1 , wherein the sequence activity model is a regression model.
13 . The method of claim 1 , wherein using the sequence activity model to identify one or more amino acid residues further comprises identifying sequences for use in a recombination-based diversity generation mechanism, wherein said sequences comprise variations in the one or more amino acid residues identified in (c).
14 . The method of claim 1 , wherein using the sequence activity model comprises identifying a sequence predicted by the model to have a highest value of the desired activity.
15 . The method of claim 1 , wherein using the sequence activity model to identify one or more amino acid residues comprises using the sequence activity model to rank residue positions in order of impact on the desired activity.
16 . The method of claim 1 , wherein using the model comprises using the model as a fitness function in a genetic algorithm.
17 . The method of claim 16 , wherein the genetic algorithm is employed to select a sequence predicted by the model to have a highest value of the desired activity.
18 . The method of claim 1 , wherein using the sequence activity model to identify one or more amino acid residues at specific positions comprises identifying one or more sequences for use in generating a new protein variant library.
19 . The method of claim 18 , wherein the one or more sequences for use in generating the new protein variant library are oligonucleotide sequences encoding variations of the one or more identified amino acid residues.
20 . The method of claim 19 , wherein the oligonucleotide sequences encode at least a portion of (i) a naturally occurring parent protein having the highest activity among naturally occurring parent proteins, or (ii) a sequence predicted by the sequence activity model to have the highest activity.
21 . The method of claim 18 , further comprising developing a new sequence activity model using activity and sequence data characterizing the new protein variant library.
22 . The method of claim 1 , wherein the one or more amino acid residues identified in (c) are identified in a reference sequence predicted using the sequence activity model or a reference sequence that describes a member of the protein variant library.
23 . The method of claim 1 , wherein the training set of a protein variant library comprises proteins that were obtained by performing DNA fragmentation-mediated recombination or a synthetic oligonucleotide-mediated recombination on nucleic acids encoding all or part of one or more naturally occurring parent proteins.
24 . A computer program product comprising a machine readable medium on which is provided program instructions for identifying amino acid residues for variation in a protein variant library in order to affect a desired activity, said instructions comprising:
(a) code for receiving data characterizing a training set of a protein variant library, wherein the data provides activity and sequence information for each protein variant in the training set; (b) code for developing, from the data, a sequence-activity model that predicts activity as a function of amino acid residue type and corresponding position in a protein sequence, wherein the sequence-activity model includes one or more non-linear terms, each representing an interaction between two or more amino acid residues in the protein sequence; and (c) code for using the sequence-activity model to identify one or more amino acid residues at specific positions for variation to impact the desired activity.
25 - 40 . (canceled)
41 . A method for identifying nucleotides for variation in nucleic acids encoding a protein variant library in order to affect a desired activity, said method comprising:
(a) receiving data characterizing a training set of a protein variant library, wherein the data provides activity and nucleotide sequence information for each protein variant in the training set; (b) from the data, developing a sequence activity model that predicts activity as a function of nucleotide types and corresponding position in the nucleotide sequence, wherein the sequence-activity model includes one or more non-linear terms, each representing an interaction between two or more amino acid residues in the protein sequence; and (c) using the sequence activity model to rank positions in a nucleotide sequence and/or nucleotide types at specific positions in the nucleotide sequence in order of impact on the desired activity; (d) using the ranking to identify one or more nucleotides, in the nucleotide sequence, that are to be varied or fixed in order to impact the desired activity.
42 . The method of claim 41 , wherein the nucleotides to be varied are codons encoding particular amino acids.
43 . The method of claim 42 , wherein at least one of the non-linear terms is a cross-product term comprising a product of one variable representing the presence of a codon encoding one interacting residue and another variable representing the presence of another codon encoding a different interacting residue.
44 . The method of claim 43 , wherein the sequence-activity model comprises a sum of said at least one cross-product term and one or more linear terms, each representing the presence of a codon encoding a variable residue in the training set.
45 . The method of claim 43 , wherein developing said sequence-activity model comprises selecting one or more cross-product terms from a group of potential cross-product terms.
46 . The method of claim 45 , wherein selecting the one or more cross-product terms comprises running a genetic algorithm to select a cross-product terms based upon the predictive ability of various models employing different cross-product terms.
47 . The method of claim 41 , wherein the activity is a function of expression of nucleic acids.
48 . A computer program product comprising a machine readable medium on which is provided program code for identifying nucleotides for variation in nucleic acids encoding a protein variant library in order to affect a desired activity, said program code comprising:
(a) code for receiving data characterizing a training set of a protein variant library, wherein the data provides activity and nucleotide sequence information for each protein variant in the training set; (b) code for developing, from said data, a sequence activity model that predicts activity as a function of nucleotide types and corresponding position in the nucleotide sequence, wherein the sequence-activity model includes one or more non-linear terms, each representing an interaction between two or more amino acid residues in the protein sequence; and (c) code for using the sequence activity model to rank positions in a nucleotide sequence and/or nucleotide types at specific positions in the nucleotide sequence in order of impact on the desired activity; (d) code for using the ranking to identify one or more nucleotides, in the nucleotide sequence, that are to be varied or fixed in order to impact the desired activity.
49 - 54 . (canceled)Join the waitlist — get patent alerts
Track US2005084907A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.