US2020040347A1PendingUtilityA1
Systems and Methods for Identifying and Expressing Gene Clusters
Assignee: UNIV LELAND STANFORD JUNIORPriority: Nov 16, 2016Filed: Nov 16, 2017Published: Feb 6, 2020
Est. expiryNov 16, 2036(~10.3 yrs left)· nominal 20-yr term from priority
G16B 30/00C12N 15/81G16B 10/00C12N 15/1079C12N 15/1089C12N 15/79G16B 25/00G16B 20/30G16B 5/00G16B 20/00G16B 30/10
59
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Methods for identifying biosynthetic gene clusters that include genes for producing compounds that interact with specific target proteins are disclosed. Some methods relate to bioinformatics methods for identifying and/or prioritizing biosynthetic gene clusters. Related systems, components, and tools for the identification and expression of such gene clusters are also disclosed.
Claims
exact text as granted — not AI-modified1 .- 9 . (canceled)
10 . A method for producing a small molecule for modulating a first target protein, the method comprising:
selecting, from a database comprising a list of biosynthetic gene clusters, one or more gene clusters that include or are positioned proximal to a region that encodes a protein that is identical with or homologous to the first target protein, expressing the gene cluster or a plurality of genes from the gene cluster in a host cell; and isolating a compound produced by the gene cluster.
11 . The method of claim 10 , wherein the one or more gene clusters are selected from the group consisting of (1) clusters that comprise one or more polyketide synthases and (2) clusters that comprise one or more non-ribosomal peptide synthetases, (3) clusters that comprise one or more terpene synthases, (4) clusters that comprises one or more UbiA-type terpene cyclases, and (5) clusters that comprise one or more dimethylallyl transferases.
12 . The method of claim 10 , wherein the protein that is encoded by the region that is included in or positioned proximal to the biosynthetic gene cluster is identical to or has greater than 30% homology to the first target protein.
13 . The method of claim 10 , wherein the region that encodes the protein that is identical with or homologous to the first target protein is within 20,000 base pairs of a region of a portion of the gene cluster that encodes a polyketide synthase, a non-ribosomal peptide synthetase, a terpene synthetase, a UbiA-type terpene cyclase, or a dimethylallyl transferase.
14 . (canceled)
15 . The method of claim 10 , wherein selecting the one or more gene clusters comprises operating a computer, wherein operation of the computer comprises running an algorithm that takes into account both an input sequence for the first target protein and sequence information from a database that includes sequence information from a plurality of species such that the computer returns information corresponding to one or more gene clusters.
16 . The method of claim 15 , wherein the algorithm takes into account the phylogenetic relationship between gene clusters in the database.
17 . The method of claim 10 , wherein the one or more gene clusters include a coding sequence for a protein that is an extracellular protein, a membrane tethered protein, a protein involved in a transport or secretion pathway, a protein homologous to a protein involved in a transport or secretion pathway, a protein with a peptide targeting signal, a protein with a terminal sequence with homology to a targeting signal, an enzyme that degrades small molecules, or a protein with homology to an enzyme that degrades small molecules.
18 . (canceled)
19 . The method of claim 18 , further comprising screening the isolated compound for modulation of an activity of the first target protein.
20 . The method of claim 18 , wherein the cluster-encoded protein that is homologous to the first target protein is resistant to modulation by the isolated compound when compared to modulation of the first target protein.
21 . The method of claim 20 , wherein the compound is not toxic to the species from which the cluster originates due to one or more of (1) sequence differences between the first target protein and the cluster-encoded protein, (2) spatial separation of the compound from the cluster-encoded protein and (3) high expression levels for the cluster-encoded protein.
22 .- 33 . (canceled)
33 . The method of claim 10 , wherein the gene cluster is a gene cluster of a non-yeast fungus.
34 . (canceled)
35 . The method of claim 10 , wherein the first target protein is a human protein.
36 .- 37 . (canceled)
38 . A system for identifying one or more biosynthetic gene clusters for introduction into a host organism to produce one or more compounds that modulate a specific target protein, the system comprising:
a processor; a memory containing a gene cluster identification application; wherein the gene cluster identification application directs the processor to:
load data describing at least one target protein into the memory;
load data describing a plurality of biosynthetic gene clusters into the memory;
score each of the plurality of biosynthetic gene clusters based upon:
performing a homolog search for each biosynthetic gene cluster to determine a presence of at least one homolog of a target protein within or adjacent the biosynthetic gene cluster;
confidence of homology of the at least one target protein to at least one gene in a biosynthetic gene cluster;
a fraction of a homologous gene that meets an identity threshold;
a total number of genes homologous to the at least one target protein present in the entire genome of an organism;
homology of the at least one homolog of at least one target protein within or adjacent the biosynthetic gene cluster to genes in the target protein's genome;
phylogenetic relationship of the at least one target protein to a gene in a cluster;
expected number of homologs of the at least one target protein in or adjacent to a biosynthetic cluster; or
a likelihood that at least one target protein is essential for cellular process in the natural environment; and
output a report identifying one or more biosynthetic gene clusters that are most likely to produce a compound that modulates the at least one target protein.
39 .- 40 . (canceled)
41 . A method for producing a compound that binds a protein of interest, the method comprising:
obtaining sequence information for a plurality of contiguous sequences, wherein each contiguous sequence includes a biosynthetic gene cluster and flanking genomic sequences;
analyzing the contiguous sequences for the presence of a gene that encodes a protein with homology to the protein of interest, and
selecting a biosynthetic gene cluster which includes, or is proximal to, a gene that encodes a protein that is homologous to the protein of interest,
expressing the biosynthetic gene cluster or a plurality of genes from the biosynthetic gene cluster in a host cell; and isolating a compound produced by the biosynthetic gene cluster.
42 . The method of claim 41 , wherein the contiguous nucleotide sequence is less than 40,000 base pairs in length.
43 .- 76 . (canceled)
77 . The method of claim 10 , wherein the first target protein is a mammalian protein.
78 . The method of claim 18 , wherein said host cell is a yeast cell.
79 . The method of claim 18 , wherein each expressed gene of the plurality of genes from the gene cluster is expressed under the control of a different promoter.Join the waitlist — get patent alerts
Track US2020040347A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.