Identification and comparison of protein-protein interactions that occur in populations and identification of inhibitors of these interactors
Abstract
Methods are described for detecting protein-protein interactions, among two populations of proteins, each having a complexity of at least 1,000. For example, proteins are fused either to the DNA-binding domain of a transcriptional activator or to the activation domain of a transcriptional activator. Two yeast strains, of the opposite mating type and carrying one type each of the fusion proteins are mated together. Productive interactions between the two halves due to protein-protein interactions lead to the reconstitution of the transcriptional activator, which in turn leads to the activation of a reporter gene containing a binding site for the DNA-binding domain. This analysis can be carried out for two or more populations of proteins. The differences in the genes encoding the proteins involved in the protein-protein interactions are characterized, thus leading to the identification of specific protein-protein interactions, and the genes encoding the interacting proteins, relevant to a particular tissue, stage or disease. Furthermore, inhibitors that interfere with these protein-protein interactions are identified by their ability to inactivate a reporter gene. The screening for such inhibitors can be in a multiplexed format where a set of inhibitors will be screened against a library of interactors.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of detecting one or more protein-protein interactions comprising
(a) recombinantly expressing within a population of host cells
(i) a first population of first fusion proteins, each said first fusion protein comprising a first protein sequence and a DNA binding domain in which the DNA binding domain is the same in each said first fusion protein, and in which said first population of first fusion proteins has a complexity of at least 1,000; and
(ii) a second population of second fusion proteins, each said second fusion protein comprising a second protein sequence and a transcriptional regulatory domain of a transcriptional regulator, in which the transcriptional regulatory domain is the same in each said second fusion protein, such that a first fusion protein is co-expressed with a second fusion protein in host cells, and wherein said host cells contain at least one nucleotide sequence operably linked to a promoter driven by one or more DNA binding sites recognized by said DNA binding domain such that interaction of a first fusion protein with a second fusion protein results in regulation of transcription of said at least one nucleotide sequence by said regulatory domain, and in which said second population of second fusion proteins has a complexity of at least 1,000; and
(b) detecting said regulation of transcription of said at least one nucleotide sequence, thereby detecting an interaction between a first fusion protein and a second fusion protein.
2 . The method according to claim 1 in which the regulatory domain is an activation domain, and said regulation of transcription is activation of transcription.
3 . The method according to claim 1 in which the first and second populations of fusion proteins are each expressed from chimeric genes comprising cDNA sequences from an uncharacterized sample of a population of cDNA from mammalian RNA.
4 . The method according to claim 2 in which the first and second populations of fusion proteins comprise first and second protein sequences, respectively, that are encoded by DNA sequences representative of the same DNA population.
5 . The method according to claim 2 in which the first and second populations of fusion proteins comprise first and second protein sequences, respectively, that are different.
6 . The method according to claim 4 in which the first and second populations of fusion proteins are each expressed from chimeric genes comprising cDNA sequences of total mammalian RNA or polyA + RNA of a cell.
7 . The method according to claim 5 in which the first and second populations of fusion proteins are each expressed from chimeric genes comprising cDNA sequences of mammalian RNA, and the first population of first fusion proteins is expressed from chimeric genes comprising cDNA sequences of diseased human tissue, and the second population of second fusion proteins is expressed from chimeric genes comprising cDNA sequences of non-diseased human tissue.
8 . The method according to claim 6 in which the cDNA sequences are of diseased human tissue.
9 . The method according to claim 2 in which said first or second population of fusion proteins has a complexity of at least 10,000.
10 . The method according to claim 2 in which said first or second population of fusion proteins has a complexity of at least 50,000.
11 . The method according to claim 2 in which said first and second populations of fusion proteins each has a complexity of at least 10,000.
12 . The method according to claim 2 in which said first and second populations of fusion proteins each has a complexity of at least 50,000.
13 . The method according to claim 1 in which said first population of first fusion proteins is expressed from a first plasmid expression vector that expresses a first selectable marker, and the second population of second fusion proteins is expressed from a second plasmid expression vector that expresses a second selectable marker, and in which the population of host cells is incubated in an environment in which substantial death of host cells occurs in the absence of expression of the first and second selectable markers.
14 . The method according to claim 2 in which the population of host cells is a population of mammalian host cells.
15 . The method according to claim 2 in which the population of host cells is a population of yeast host cells.
16 . The method according to claim 2 in which the population of host cells is a population of bacterial host cells.
17 . A method of detecting an inhibitor of a protein-protein interaction comprising
(a) incubating a population of cells, said population comprising cells recombinantly expressing a pair of interacting proteins, said pair consisting of a first protein and a second protein, in the presence of one or more candidate molecules among which it is desired to identify an inhibitor of the interaction between said first protein and said second protein, in an environment in which substantial death of said cells occurs (i) when said first protein and second protein interact, or (ii) if said cells lack a recombinant nucleic acid encoding said first protein or a recombinant nucleic acid encoding said second protein; and (b) detecting those cells that survive said incubating step, thereby detecting the presence of an inhibitor of said interaction in said cells.
18 . The method according to claim 17 in which the cells are yeast cells.
19 . The method according to claim 17 in which the first protein and the second protein are first and second fusion proteins, respectively, between which an interaction is detected according to the method of claim 1 .
20 . The method according to claim 17 in which the first protein and the second protein are first and second fusion proteins, respectively, between which an interaction is detected according to the method of claim 2 .
21 . A method of detecting one or more protein-protein interactions comprising
(a) recombinantly expressing in a first population of yeast cells of a first mating type, a first population of first fusion proteins, each first fusion protein comprising a first protein sequence and a DNA binding domain, in which the DNA binding domain is the same in each said first fusion protein; wherein said first population of yeast cells contains a first nucleotide sequence operably linked to a promoter driven by one or more DNA binding sites recognized by said DNA binding domain such that an interaction of a first fusion protein with a second fusion protein, said second fusion protein comprising a transcriptional activation domain, results in increased transcription of said first nucleotide sequence, and in which said first population of first fusion proteins has a complexity of at least 1,000; (b) negatively selecting to eliminate those yeast cells expressing said first population of first fusion proteins in which said increased transcription of said first nucleotide sequence occurs in the absence of said second fusion protein; (c) recombinantly expressing in a second population of yeast cells of a second mating type different from said first mating type, a second population of said second fusion proteins, each second fusion protein comprising a second protein sequence and an activation domain of a transcriptional activator, in which the activation domain is the same in each said second fusion protein, and in which said second population of second fusion proteins has a complexity of at least 1,000; (d) mating said first population of yeast cells with said second population of yeast cells to form a population of diploid yeast cells, wherein said population of diploid yeast cells contains a second nucleotide sequence operably linked to a promoter driven by a DNA binding site recognized by said DNA binding domain such that an interaction of a first fusion protein with a second fusion protein results in increased transcription of said second nucleotide sequence, in which the first and second nucleotide sequences can be the same or different; and (e) detecting said increased transcription of said first and/or second nucleotide sequence, thereby detecting an interaction between a first fusion protein and a second fusion protein.
22 . The method according to claim 21 in which said negatively selecting is carried out by a method comprising incubating said first population of yeast cells expressing said first population of first fusion proteins in an environment in which substantial death of said first population of host cells occurs if said increased transcription occurs.
23 . The method according to claim 22 in which said first nucleotide sequence comprises a functional URA3 coding sequence, and said environment contains 5-fluoroorotic acid.
24 . The method according to claim 22 in which said first nucleotide sequence comprises a functional LYS2 coding sequence, and said environment comprises α-amino-adipate.
25 . The method according to claim 21 in which said second nucleotide sequence is different from said first nucleotide sequence.
26 . The method according to claim 23 in which said second nucleotide sequence comprises a functional lacZ coding sequence.
27 . The method according to claim 21 in which said first and second nucleotide sequences are selected from the group consisting of the functional coding sequences of URA3, HIS3, lacZ, GFP, LEU2, LYS2, ADE2, TRP1, CAN1, CYH2, GUS, CUP1, and CAT.
28 . The method according to claim 21 in which said DNA binding domain is selected from the group consisting of the DNA binding domains of GAL4, GCN4, ARD1, LEX A, and Ace1N.
29 . The method according to claim 21 in which said transcription activation domain is selected from the group consisting of the activation domains of GAL4, GCN4, ARD1, herpes simplex virus VP16, and Ace1C.
30 . The method according to claim 21 in which said first population of yeast cells and said second population of yeast cells do not contain functional counterparts of said first and second nucleotide sequences that are not operably linked to a promoter driven by said one or more DNA binding sites.
31 . The method according to claim 21 in which said DNA binding domain is a GAL4 or LEX A DNA binding domain, and said transcription activation domain is a GAL4 or herpes simplex virus VP16 activation domain.
32 . The method according to claim 21 in which the first and second populations of fusion proteins are each expressed from chimeric genes comprising cDNA sequences from an uncharacterized sample of a population of cDNA from mammalian RNA.
33 . The method according to claim 21 in which the first and second populations of fusion proteins comprise first and second protein sequences, respectively, that are encoded by DNA sequences representative of the same DNA population.
34 . The method according to claim 21 in which the first and second populations of fusion proteins comprise first and second protein sequences, respectively, that are different.
35 . The method according to claim 33 in which the first and second populations of fusion proteins are each expressed from chimeric genes comprising cDNA sequences of mammalian RNA.
36 . The method according to claim 34 in which the first and second populations of fusion proteins are each expressed from chimeric genes comprising cDNA sequences of mammalian RNA, and the first population of first fusion proteins is expressed from chimeric genes comprising cDNA sequences of diseased human tissue, and the second population of second fusion proteins is expressed from chimeric genes comprising cDNA sequences of non-diseased human tissue.
37 . The method according to claim 35 in which the cDNA sequences are of diseased human tissue.
38 . The method according to claim 21 in which said first or second population of fusion proteins has a complexity of at least 10,000.
39 . The method according to claim 21 in which said first or second population of fusion proteins has a complexity of at least 50,000.
40 . The method according to claim 21 in which said first and second populations of fusion proteins each has a complexity of at least 10,000.
41 . The method according to claim 21 in which said first and second populations of fusion proteins each has a complexity of at least 50,000.
42 . The method according to claim 21 , 22 or 23 in which in said mating step at least 5.8×10 8 matings are done.
43 . The method according to claim 21 , 22 or 23 in which in said mating step at least 8.5×10 9 matings are done.
44 . The method according to claim 21 , 22 or 23 in which in said mating step at least 8.5×10 10 matings are done.
45 . The method according to claim 21 , 22 or 23 in which said mating is performed on solid medium.
46 . The method according to claim 21 in which the first population of first fusion proteins is expressed from a first plasmid expression vector that expresses a first selectable marker, and the second population of second fusion proteins is expressed from a second plasmid expression vector that expresses a second selectable marker, and in which the first population of yeast cells is incubated in a first environment in which substantial death of yeast cells occurs in the absence of expression of the first selectable marker, and the second population of yeast cells is incubated in a second environment in which substantial death of yeast cells occurs in the absence of expression of the second selectable marker.
47 . The method according to claim 21 in which the yeast cells are Saccharomyces cerevisiae.
48 . The method according to claim 17 in which the cells are yeast cells, and in which the first protein and the second protein are first and second fusion proteins, respectively, between which an interaction is detected according to the method of claim 21 .
49 . The method according to claim 48 in which the yeast cells contain functional URA3 coding sequences under the control of a promoter driven by a DNA binding site recognized by said DNA binding domain of said first fusion protein, and said environment contains 5-fluoroorotic acid.
50 . The method according to claim 48 or 49 in which said one or more candidate molecules are provided to said at least one diploid yeast cell by introducing into said diploid yeast cell one or more recombinant nucleic acids encoding said one or more candidate molecules, such that said one or more candidate molecules are expressed within said diploid yeast cell.
51 . The method according to claim 48 or 49 in which said one or more candidate molecules are provided to said at least one diploid yeast cell by incubating said diploid yeast cell in an environment comprising said one or more candidate molecules.
52 . A method of detecting one or more protein-protein interactions present within a first protein population and absent within a second protein population comprising
(a) carrying out the method of claim 1 wherein said first protein sequences of said first fusion proteins and said second protein sequences of said second fusion proteins are encoded by DNA sequences representative of the same first DNA population, thereby detecting one or more protein-protein interactions; (b) carrying out the method of claim 1 wherein said first protein sequences of said first fusion proteins and said second protein sequences of said second fusion proteins are encoded by DNA sequences representative of the same second DNA population, said second DNA population differing from said first DNA population, thereby detecting one or more protein-protein interactions; and (c) comparing the one or more protein-protein interactions detected in step (a) with the one or more protein-protein interactions detected in step (b).
53 . A method of detecting one or more protein-protein interactions present within a first protein population and absent within a second protein population comprising
(a) carrying out the method of claim 21 wherein said first protein sequences of said first fusion proteins and said second protein sequences of said second fusion proteins are encoded by DNA sequences representative of the same first DNA population, thereby detecting one or more protein-protein interactions; (b) carrying out the method of claim 21 wherein said first protein and said second protein sequences of said second fusion proteins are encoded by DNA sequences representative of the same second DNA population, said second DNA population differing from said first DNA population, thereby detecting one or more protein-protein interactions; and (c) comparing the one or more protein-protein interactions detected in step (a) with the one or more protein-protein interactions detected in step (b).
54 . A method of detecting one or more protein-protein interactions comprising
(a) introducing into a first population of cells of Saccharomyces cerevisiae a first population of first plasmids, each said first plasmid encoding and capable of expressing in the first population of cells (i) TRP1, and (ii) a first population of first fusion proteins, each said first fusion protein comprising a GAL4 DNA binding domain and a first protein sequence, in which said first population of first fusion proteins has a complexity of at least 10,000, and in which said first population of cells (i) is of a first mating type selected from the group consisting of a and α, (ii) is mutant in endogenous URA3 and HIS3, (iii) contains functional URA3 coding sequences under the control of a promoter containing GAL4 binding sites, and (iv) contains functional lacZ coding sequences under the control of a promoter containing GAL4 binding sites; (b) introducing into a second population of cells of Saccharomyces cerevisiae a second population of second plasmids, each said second plasmids encoding and capable of expressing in the second population of cells (i) LEU2, and (ii) a second population of second fusion proteins, each said second fusion protein comprising a GAL4 transcriptional activation domain and a second protein sequence, in which said second population of second fusion proteins has a complexity of at least 10,000, and in which said second population of cells (i) is of a second mating type different from said first mating type and selected from the group consisting of a and α, (ii) is mutant in endogenous URA3 and HIS3, (iii) contains functional HIS3 coding sequences under the control of a promoter containing GAL4 binding sites, and (iv) contains functional lacZ coding sequences under the control of a promoter containing GAL4 binding sites; (c) after step (a), incubating said first population of cells in an environment lacking tryptophan and containing 5-fluoroorotic acid; (d) pooling surviving cells from said first population after step (c); (e) after step (b), incubating said second population of cells in an environment lacking leucine; (f) pooling surviving cells from said second population after step (e); (g) mating the pooled cells from said first population and the pooled cells from said second population by mixing the cells together, applying the cells to a solid medium and incubating the cells, to form diploid cells; and (h) incubating the diploid cells in an environment lacking uracil, histidine, tryptophan and leucine, to select diploid cells containing a said first plasmid and a said second plasmid and in which transcription of the URA3 and HIS3 coding sequences has been activated, thereby indicating that a first fusion protein has interacted with a second fusion protein within the diploid cell, thereby detecting one or more protein-protein interactions.
55 . The method according to claim 54 in which the pooled cells from said first population and the pooled cells from said second population in said mating step are each at least 8.5×10 9 in number.
56 . The method according to claim 54 in which the first and second populations of fusion proteins are each expressed from chimeric genes comprising cDNA sequences of mammalian RNA.
57 . The method according to claim 1 which further comprises obtaining a purified DNA encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a host cell in which said regulation of transcription is detected.
58 . The method according to claim 1 which further comprises obtaining a purified DNA encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a host cell in which said regulation of transcription is detected.
59 . The method according to claim 57 which further comprises obtaining a purified DNA encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a host cell in which said regulation of transcription is detected.
60 . The method according to claim 21 which further comprises obtaining a purified DNA encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a yeast cell in which said increased transcription is detected in step (e).
61 . The method according to claim 21 which further comprises obtaining a purified DNA encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a yeast cell in which said increased transcription is detected in step (e).
62 . The method according to claim 60 which further comprises obtaining a purified DNA encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a yeast cell in which said increased transcription is detected in step (e).
63 . The method according to claim 54 which further comprises obtaining a purified DNA encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a cell which survives said incubating step (h) and is thereby selected.
64 . The method according to claim 54 which further comprises obtaining a purified DNA encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a cell which survives said incubating step (h) and is thereby selected.
65 . The method according to claim 63 which further comprises obtaining a purified DNA encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a cell which survives said incubating step (h) and is thereby selected.
66 . The method according to claim 1 which further comprises amplifying a DNA fragment encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a host cell in which said regulation of transcription is detected.
67 . The method according to claim 1 which further comprises amplifying a DNA fragment encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a host cell in which said regulation of transcription is detected.
68 . The method according to claim 1 which further comprises:
(c) amplifying a DNA fragment encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a host cell in which said regulation of transcription is detected; and
(d) amplifying a DNA fragment encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from said host cell.
69 . The method according to claim 21 which further comprises amplifying a DNA fragment encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence from a yeast cell in which said increased transcription is detected in step (e).
70 . The method according to claim 21 which further comprises amplifying a DNA fragment encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a yeast cell in which said increased transcription is detected in step (e).
71 . The method according to claim 21 which further comprises:
(f) amplifying a DNA fragment encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a yeast cell in which said increased transcription is detected in step (e); and
(g) amplifying a DNA fragment encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from said yeast cell.
72 . The method according to claim 46 in which said first and second plasmid expression vectors are replicable both in yeast cells and in E. coli.
73 . The method according to claim 72 which further comprises obtaining a first purified DNA encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a yeast cell in which said increased transcription is detected in step (e).
74 . The method according to claim 72 which further comprises obtaining a purified DNA encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a yeast cell in which said increased transcription is detected in step (e).
75 . The method according to claim 73 which further comprises obtaining a second purified DNA encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a yeast cell in which said increased transcription is detected in step (e).
76 . The method according to claim 73 in which said first purified DNA is obtained by recovering said first plasmid expression vector from said yeast cell in which said increased transcription is detected, transforming an E. coli with said first plasmid expression vector, incubating said E. coli , and recovering said first plasmid expression vector from the incubated E. coli.
77 . The method according to claim 74 in which said purified DNA is obtained by recovering said second plasmid expression vector from said yeast cell in which said increased transcription is detected, transforming an E. coli with said second plasmid expression vector, incubating said E. coli , and recovering said second plasmid expression vector from the incubated E. coli.
78 . The method according to claim 75 in which first purified DNA is obtained by recovering said first plasmid expression vector from said yeast cell in which said increased transcription is detected, transforming an E. coli with said first plasmid expression vector, incubating said E. coli , and recovering said first plasmid expression vector from the incubated E. coli ; and in which second purified DNA is obtained by recovering said second plasmid expression vector from said yeast cell in which said increased transcription is detected, transforming an E. coli with said second plasmid expression vector, incubating said E. coli , and recovering said second plasmid, expression vector from the incubated E. coli.
79 . The method according to claim 76 which further comprises sequencing at least a portion of said first purified DNA to determine the sequence of said first protein sequence.
80 . The method according to claim 77 which further comprises sequencing at least a portion of said purified DNA to determine the sequence of said second protein sequence.
81 . The method according to claim 54 which further comprises amplifying a DNA fragment encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a cell which survives said incubating step (h) and is thereby selected.
82 . The method according to claim 54 which further comprises amplifying a DNA fragment encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a cell which survives said incubating step (h) and is thereby selected.
83 . The method according to claim 54 which further comprises:
(i) amplifying a DNA fragment encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a cell which survives said incubating step (h) and is thereby selected; and
(j) amplifying a DNA fragment encoding said second fusion protein or encoding portion thereof comprising said second protein sequence, from said cell.
84 . The method according to claim 54 in which said first plasmids and said second plasmids are replicable both in yeast cells and in E. coli.
85 . The method according to claim 84 which further comprises obtaining a first purified DNA encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a yeast cell selected in step (h).
86 . The method according to claim 84 which further comprises obtaining a purified DNA encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a yeast cell selected in step (h).
87 . The method according to claim 85 which further comprises obtaining a second purified DNA encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a yeast cell selected in step (h).
88 . The method according to claim 85 in which said first purified DNA is obtained by recovering said first plasmid from said yeast cell selected in step (h), transforming an E. coli with said first plasmid, incubating said E. coli , and recovering said first plasmid from the incubated E. coli.
89 . The method according to claim 86 in which purified DNA is obtained by recovering said second plasmid from said yeast cell selected in step (h), transforming an E. coli with said second plasmid, incubating said E. coli , and recovering said second plasmid from the incubated E. coli.
90 . The method according to claim 78 in which said first purified DNA is obtained by recovering said first plasmid from said yeast cell selected in step (h), transforming an E. coli with said first plasmid, incubating said E. coli , and recovering said first plasmid from the incubated E. coli ; and in which said second purified DNA is obtained by recovering said second plasmid from said yeast cell selected in step (h), transforming an E. coli with said second plasmid, incubating said E. coli , and recovering said second plasmid from the incubated E. coli.
91 . The method according to claim 88 which further comprises sequencing at least a portion of said first purified DNA to determine the sequence of said first protein sequence.
92 . The method according to claim 89 which further comprises sequencing at least a portion of said purified DNA to determine the sequence of said second protein sequence.
93 . The method according to claim 66 , 69 or 81 which further comprises sequencing at least a portion of said purified DNA to determine the sequence of said first protein sequence.
94 . The method according to claim 67 , 70 or 82 which further comprises sequencing at least a portion of said purified DNA to determine the sequence of said second protein sequence.
95 . The method according to claim 66 in which said amplifying is carried out of DNA fragments from a plurality of host cells in which said regulation of transcription is detected, and a sample comprising said resulting amplified DNA fragments is subjected to a method for identifying, classifying, or quantifying one or more nucleic acids in the sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal, whereby said one or more nucleic acids in said sample are identified, classified, or quantified.
96 . The method according to claim 69 in which said amplifying is carried out of DNA fragments from a plurality of yeast cells in which said increased transcription is detected in step (e), and a sample comprising said resulting amplified DNA fragments is subjected to a method for identifying, classifying, or quantifying one or more nucleic acids in the sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal,
whereby said one or more nucleic acids in said sample are identified, classified, or quantified.
97 . The method according to claim 81 in which said amplifying is carried out of DNA fragments from a plurality of cells which survive said incubating step (h), and a sample comprising said resulting amplified fragments is subjected to a method for identifying, classifying, or quantifying one or more nucleic acids in the sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal,
whereby said one or more nucleic acids in said sample are identified, classified, or quantified.
98 . The method according to claim 67 in which said amplifying is carried out of DNA fragments from a plurality of host cells in which said regulation of transcription is detected, and a sample comprising said resulting amplified DNA fragments is subjected to a method for identifying, classifying, or quantifying one or more nucleic acids in the sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal,
whereby said one or more nucleic acids fin said sample are identified, classified, or quantified.
99 . The method according to claim 70 in which said amplifying is carried out of DNA fragment from a plurality of yeast cells in which said increased transcription is detected in step (e), and a sample comprising said resulting amplified DNA fragments is subjected to a method for identifying, classifying, or quantifying one or more nucleic acids in the sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal,
whereby said one or more nucleic acids in said sample are identified, classified, or quantified.
100 . The method according to claim 82 in which said amplifying is carried out of DNA fragments from a plurality of cells which survive said incubating step (h), and a sample comprising said resulting amplified fragments is subjected to a method for identifying, classifying, or quantifying one or more nucleic acids in the sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal,
whereby said one or more nucleic acids in said sample are identified, classified, or quantified.
101 . The method according to claim 66 in which said amplifying is carried out of DNA fragments from a plurality of host cells in which said regulation of transcription is detected, and a sample comprising said resulting amplified DNA fragments is subjected to a method for identifying, classifying, or quantifying DNA molecules in the sample, said method comprising:
(a) digesting said sample with one or more restriction endonucleases, each said restriction endonuclease recognizing a subsequence recognition site and digesting DNA at said recognition site to produce fragments with 5′ overhangs;
(b) contacting said produced fragments with shorter and longer oligodeoxynucleotides, each said shorter oligodeoxynucleotide hybridizable with a said 5′ overhang and having no terminal phosphates, each said longer oligodeoxynucleotide hybridizable with a said shorter oligodeoxynucleotide;
(c) ligating said longer oligodeoxynucleotides to said 5′ overhangs on said fragments to produce ligated DNA fragments;
(d) extending said ligated DNA fragments by synthesis with a DNA polymerase to produce blunt-ended double stranded DNA fragments;
(e) amplifying said blunt-ended double stranded DNA fragments by a method comprising contacting said blunt-ended double stranded DNA fragments with a DNA polymerase and primer oligodeoxynucleotides, each said primer oligodeoxynucleotide having a sequence comprising that of one of the longer oligodeoxynucleotides;
(f) determining the length of the amplified DNA fragments produced in step (e); and
(g) searching a DNA sequence database, said database comprising a plurality of known DNA sequences that may be present in the sample, for sequences matching one or more of said fragments of determined length, a sequence from said database matching a fragment of determined length when the sequence from said database comprises recognition sites of said one or more restriction endonucleases spaced apart by the determined length, whereby DNA molecules in said sample are identified, classified, or quantified.
102 . The method according to claim 69 in which said amplifying is carried out of DNA fragments from a plurality of yeast cells in which said increased transcription is detected in step (e), and a sample comprising said resulting amplified DNA fragments is subjected to a method for identifying, classifying, or quantifying DNA molecules in the sample, said method comprising:
(a) digesting said sample with one or more restriction endonucleases, each said restriction endonuclease recognizing a subsequence recognition site and digesting DNA at said recognition site to produce fragments with 5′ overhangs;
(b) contacting said produced fragments with shorter and longer oligodeoxynucleotides, each said shorter oligodeoxynucleotide hybridizable with a said 5′ overhang and having no terminal phosphates, each said longer oligodeoxynucleotide hybridizable with a said shorter oligodeoxynucleotide;
(c) ligating said longer oligodeoxynucleotides to said 5′ overhangs on said fragments to produce ligated DNA fragments;
(d) extending said ligated DNA fragments by synthesis with a DNA polymerase to produce blunt-ended double stranded DNA fragments;
(e) amplifying said blunt-ended double stranded DNA fragments by a method comprising contacting said blunt-ended double stranded DNA fragments produced in step (d) with a DNA polymerase and primer oligodeoxynucleotides, each said primer oligodeoxynucleotide having a sequence comprising that of one of the longer oligodeoxynucleotides;
(f) determining the length of the amplified DNA fragments produced in step (e); and
(g) searching a DNA sequence database, said database comprising a plurality of known DNA sequences that may be present in the sample, for sequences matching one or more of said fragments of determined length, a sequence from said database matching a fragment of determined length when the sequence from said database comprises recognition sites of said one or more restriction endonucleases spaced apart by the determined length,
whereby DNA molecules in said sample are identified, classified, or quantified.
103 . The method according to claim 81 in which said amplifying is carried out of DNA fragments from a plurality of cells which survive said incubating step (h), and a sample comprising said resulting amplified fragments is subjected to a method for identifying, classifying, or quantifying DNA molecules in the sample, said method comprising:
(a) digesting said sample with one or more restriction endonucleases, each said restriction endonuclease recognizing a subsequence recognition site and digesting DNA at said recognition site to produce fragments with 5′ overhangs;
(b) contacting said produced fragments with shorter and longer oligodeoxynucleotides, each said shorter oligodeoxynucleotide hybridizable with a said 5′ overhang and having no terminal phosphates, each said longer oligodeoxynucleotide hybridizable with a said shorter oligodeoxynucleotide;
(c) ligating said longer oligodeoxynucleotides to said 5′ overhangs on said DNA fragments to produce ligated DNA fragments;
(d) extending said ligated DNA fragments by synthesis with a DNA polymerase to produce blunt-ended double stranded DNA fragments;
(e) amplifying said blunt-ended double stranded DNA fragments by a method comprising contacting said blunt-ended double stranded DNA fragments with a DNA polymerase and primer oligodeoxynucleotides, each said primer oligodeoxynucleotide having a sequence comprising that of one of the longer oligodeoxynucleotides;
(f) determining the length of the amplified DNA fragments produced in step (e); and
(g) searching a DNA sequence database, said database comprising a plurality of known DNA sequences that may be present in the sample, for sequences matching one or more of said fragments of determined length, a sequence from said database matching a fragment of determined length when the sequence from said database comprises recognition sites of said one or more restriction endonucleases spaced apart by the determined length,
whereby DNA molecules in said sample are identified, classified, or quantified.
104 . The method according to claim 67 in which said amplifying is carried out of DNA fragments from a plurality of host cells in which said regulation of transcription is detected, and a sample comprising said resulting amplified DNA fragments is subjected to a method for identifying, classifying, or quantifying DNA molecules in the sample, said method comprising:
(a) digesting said sample with one or more restriction endonucleases, each said restriction endonuclease recognizing a subsequence recognition site and digesting DNA at said recognition site to produce fragments with 5′ overhangs;
(b) contacting said produced fragments with shorter and longer oligodeoxynucleotides, each said shorter oligodeoxynucleotide hybridizable with a said 5′ overhang and having no terminal phosphates, each said longer oligodeoxynucleotide hybridizable with a said shorter oligodeoxynucleotide;
(c) ligating said longer oligodeoxynucleotides to said 5′ overhangs on said fragments to produce ligated DNA fragments;
(d) extending said ligated DNA fragments by synthesis with a DNA polymerase to produce blunt-ended double stranded DNA fragments;
(e) amplifying said blunt-ended double stranded DNA fragments by a method comprising contacting said blunt-ended double stranded DNA fragments with a DNA polymerase and primer oligodeoxynucleotides, each said primer oligodeoxynucleotide having a sequence comprising that of one of the longer oligodeoxynucleotides;
(f) determining the length of the amplified DNA fragments produced in step (e); and
(g) searching a DNA sequence database, said database comprising a plurality of known DNA sequences that may be present in the sample, for sequences matching one or more of said fragments of determined length, a sequence from said database matching a fragment of determined length when the sequence from said database comprises recognition sites of said one or more restriction endonucleases spaced apart by the determined length,
whereby DNA molecules in said sample are identified, classified, or quantified.
105 . The method according to claim 70 in which said amplifying is carried out of DNA fragments from a plurality of yeast cells in which said increased transcription is detected in step (e), and a sample comprising said resulting amplified DNA fragment is subjected to a method for identifying, classifying, or quantifying DNA molecules in the sample, said method comprising:
(a) digesting said sample with one or more restriction endonucleases, each said restriction endonuclease recognizing a subsequence recognition site and digesting DNA at said recognition site to produce fragments with 5′ overhangs;
(b) contacting said produced fragments with shorter and longer oligodeoxynucleotides, each said shorter oligodeoxynucleotide hybridizable with a said 5′ overhang and having no terminal phosphates, each said longer oligodeoxynucleotide hybridizable with a said shorter oligodeoxynucleotide;
(c) ligating said longer oligodeoxynucleotides to said 5′ overhangs on said fragments to produce ligated DNA fragments;
(d) extending said ligated DNA fragments by synthesis with a DNA polymerase to produce blunt-ended double stranded DNA fragments;
(e) amplifying said blunt-ended double stranded DNA fragments by a method comprising contacting said blunit-ended double stranded DNA fragments with a DNA polymerase and primer oligodeoxynucleotides, each said primer ollgodeoxynucleotide having a sequence comprising that of one of the longer oligodeoxynucleotides;
(f) determining the length of the amplified DNA fragments produced in step (e); and
(g) searching a DNA sequence database, said database comprising a plurality of known DNA sequences that may be present in the sample, for sequences matching one or more of said fragments of determined length, a sequence from said database matching a fragment of determined length when the sequence from said database comprises recognition sites of said one or more restriction endonucleases spaced apart by the determined length,
whereby DNA molecules in said sample are identified, classified, or quantified.
106 . The method according to claim 82 in which said amplifying is carried out of DNA fragments from a plurality of cells which survive said incubating step (h), and a sample comprising said resulting amplified fragments is subjected to a method for identifying, classifying, or quantifying DNA molecules in the sample, said method comprising:
(a) digesting said sample with one or more restriction endonucleases, each said restriction endonuclease recognizing a subsequence recognition site and digesting DNA at said recognition site to produce fragments with 5′ overhangs;
(b) contacting said produced fragments with shorter and longer oligodeoxynucleotides, each said shorter oligodeoxynucleotide hybridizable with a said 5′ overhang and having no terminal phosphates, each said longer oligodeoxynucleotide hybridizable with a said shorter oligodeoxynucleotide;
(c) ligating said longer oligodeoxynucleotides to said 5′ overhangs on said DNA fragments to produce ligated DNA fragments;
(d) extending said ligated DNA fragments by synthesis with a DNA polymerase to produce blunt-ended double stranded DNA fragments;
(e) amplifying said blunt-ended double stranded DNA fragments by a method comprising contacting said blunt-ended double stranded DNA fragments with a DNA polymerase and primer oligodeoxynucleotides, each said primer oligodeoxynucleotide having a sequence comprising that of one of the longer oligodeoxynucleotides;
(f) determining the length of the amplified DNA fragments produced in step (e); and
(g) searching a DNA sequence database, said database comprising a plurality of known DNA sequences that may be present in the sample, for sequences matching one or more of said fragments of determined length, a sequence from said database matching a fragment of determined length when the sequence from said database comprises recognition sites of said one or more restriction endonucleases spaced apart by the determined length,
whereby DNA molecules in said sample are identified, classified, or quantified.
107 . The method according to claim 21 which further comprises:
(f) amplifying a DNA fragment encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a plurality of yeast cells in which said increased transcription is detected in step (e);
(g) amplifying a DNA fragment encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a plurality of yeast cells in which said increased transcription is detected in step (e);
(h) pooling the amplified DNA fragments resulting from steps (f) and (g); and
(i) subjecting a sample comprising said pooled amplified DNA fragments to a method for identifying, classifying or quantifying one or more nucleic acids in a sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal.
108 . Tlhe method according to claim 54 which further comprises:
(i) amplifying a DNA fragment encoding said first fusion protein or encoding a portion thereof comprising said first protein sequence, from a cell which survives said incubating step (h) and is thereby selected;
(j) amplifying a DNA fragment encoding said second fusion protein or encoding a portion thereof comprising said second protein sequence, from a cell which survives said incubating step (h) and is thereby selected;
(k) pooling the amplified DNA fragments resulting from steps (i) and (j); and
(l) subjecting a sample comprising said pooled amplified DNA fragments to a method for identifying, classifying or quantifying one or more nucleic acids in a sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal.
109 . A method of determining one or more characteristics of or the identities of nucleic acids encoding an interacting pair of proteins from among a population of cells containing a multiplicity of different nucleic acids encoding different pairs of interacting proteins, said method comprising:
(a) designating each group of cells containing nucleic acids encoding an identical pair of interacting proteins as one point of a multidimensional array in which the intersection of axes in each dimension uniquely identifies a single said group; (b) pooling all groups along a simple axis to form a plurality of pooled groups; (c) amplifying from a first aliquot of each pooled group a plurality of first nucleic acids, each first nucleic acid comprising a sequence encoding a first protein that is one-half of a pair of interacting proteins; (d) amplifying from a second aliquot of each pooled group a plurality of second nucleic acids, each second nucleic acid comprising a sequence encoding a second protein that is the other half of the pair of interacting proteins; (e) subjecting said first nucleic acids from each pooled group to size separation; (f) subjecting said second nucleic acids from each pooled group to size separation; (g) identifying which at least one of said first nucleic acids are present in samples of first nucleic acids from a pooled group from each axes in each dimension, thereby indicating that said at least one first nucleic acid is present in said array in the group designated at the intersection of said axes in each dimension; and (h) identifying which at least one of said second nucleic acids are present in samples of a second nucleic acid from a pooled group from axes in each dimension, thereby indicating that the said at least one second nucleic acid is present in said array in the group designated at the intersection of said axes in each dimension; in which the first and second nucleic acids that are indicated to be present in said array in a group designated at the same intersection are indicated to encode interacting proteins.
110 . A method of determining one or more characteristics of or the identities of nucleic acids encoding an interacting pair of proteins from among a plurality of yeast cell colonies, each colony containing nucleic acids encoding a different pair of interacting proteins, said method comprising carrying out the method of claim 21 in which an interaction between a first fusion protein and a second fusion protein is detected in a plurality of colonies of diploid yeast cells, and which method further comprises:
(f) designating each colony in which an interaction between a first fusion protein and a second fusion protein is detected as one point of a multidimensional array in which the intersection of axes in each dimension uniquely identifies a single said colony;
(g) pooling all colonies along a simple axis to form a plurality of pooled colonies;
(h) amplifying from a first aliquot of each pooled colony a plurality of first nucleic acids, each first nucleic acid comprising a sequence encoding said first fusion protein or a portion thereof comprising said first protein sequence;
(i) amplifying from a second aliquot of each pooled colony a plurality of second nucleic acids, each second nucleic acid comprising a sequence encoding said second fusion protein or a portion thereof comprising said second protein sequence;
(j) subjecting said first nucleic acids from each pooled colony to size separation;
(k) subjecting said second nucleic acids from each pooled colony to size separation;
(l) identifying which at least one of said first nucleic acids are present in samples of first nucleic acids from a pooled colony from axes in each dimension, thereby indicating that said at least one first nucleic acid is present in said array in the colony designated at the intersection of said axes in each dimension;
(m) identifying which at least one of said second nucleic acids are present in samples of a second nucleic acid from a pooled colony from axes in each dimension, thereby indicating that the said at least one second nucleic acid is present in said array in the colony designated at the intersection of said axes in each dimension;
in which the first and second nucleic acids that are indicated to be present in said array in a colony designated at the same intersection are indicated to encode interacting protein sequences.
111 . A method of determining one or more characteristics of or the identities of DNA molecules encoding an interacting pair of proteins from among a plurality of yeast cell colonies, each colony containing DNA molecules encoding a different pair of interacting proteins, comprising carrying out the method of claim 54 in which an interaction between a first fusion protein and a second fusion protein is detected in a plurality of colonies of diploid yeast cells, and which method further comprises:
(f) designating each colony in which an interaction between a first fusion protein and a second fusion protein is detected as one point of a multidimensional array in which the intersection of axes in each dimension uniquely identifies a single said colony;
(g) pooling all colonies along a simple axis to form a plurality of pooled colonies;
(h) amplifying from a first aliquot of each pooled colony a plurality of first DNA molecules, each first DNA molecule comprising a sequence encoding said first fusion protein or a portion thereof comprising said first protein sequence;
(i) amplifying from a second aliquot of each pooled colony a plurality of second DNA molecules, each second DNA molecule comprising a sequence encoding said second fusion protein or a portion thereof comprising said second protein sequence;
(j) subjecting said first DNA molecules from each pooled colony to size separation;
(k) subjecting said second DNA molecules from each pooled colony to size separation;
(l) identifying which at least one of said first DNA molecules are present in samples of first DNA molecules from a pooled colony from axes in each dimension, thereby indicating that said at least one first DNA molecule is present in said array in the colony designated at the intersection of said axes in each dimension;
(m) identifying which at least one of said second DNA molecules are present in samples of a second DNA molecule from a pooled colony from axes in each dimension, thereby indicating that the said at least one second DNA molecule is present in said array in the colony designated at the intersection of said axes in each dimension;
in which the first and second DNA molecules that are indicated to be present in said array in a colony designated at the same intersection are indicated to encode interacting protein sequences.
112 . The method according to claim 110 or 111 in which said amplifying is by use of polymerase chain reaction.
113 . The method according to claim 110 which further comprises determining the nucleotide sequence of at least one first nucleic acid or second nucleic acid.
114 . The method according to claim 111 which further comprises determining the nucleotide sequence of at least one first DNA molecule or second DNA molecule.
115 . The method according to claim 110 which further comprises subjecting said pooled colonies of first nucleic acids to a method for identifying, classifying, or quantifying one or more nucleic acids in a sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal,
whereby said one or more nucleic acids in said sample are identified, classified, or quantified.
116 . The method according to claim 111 which further comprises subjecting said pooled colonies of first DNA molecules to a method for identifying, classifying, or quantifying one or more DNA molecules in a sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal,
whereby said one or more nucleic acids in said sample are identified, classified, or quantified.
117 . The method according to claim 110 which further comprises subjecting said pooled colonies of second nucleic acids to a method comprising a method for identifying, classifying, or quantifying one or more nucleic acids in a sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal,
whereby said one or more nucleic acids in said sample are identified, classified, or quantified.
118 . The method according to claim 111 which further comprises subjecting said pooled colonies of second DNA molecules to a method comprising a method for identifying, classifying, or quantifying one or more DNA molecules in a sample, said method comprising:
(a) probing said sample with one or more recognition means, each recognition means causing recognition of a target nucleotide subsequence or a set of target nucleotide subsequences;
(b) generating one or more signals from said sample probed by said recognition means, each generated signal arising from a nucleic acid in said sample and comprising a representation of (i) the identities of effective subsequences, each said effective subsequence being a subsequence comprising a target subsequence, or the identities of sets of effective subsequences, each said set having member effective subsequences each of which comprises a different target subsequence from one of said sets of target sequences, and (ii) the length between occurrences of effective subsequences in said nucleic acid or between one occurrence of one effective subsequence and the end of said nucleic acid; and
(c) searching a nucleotide sequence database to determine sequences that match or the absence of any sequences that match said one or more generated signals, said database comprising a plurality of known nucleotide sequences of nucleic acids that may be present in the sample, a sequence from said database matching a generated signal when the sequence from said database has both (i) the same length between occurrences of effective subsequences or the same length between one occurrence of one effective target subsequence and the end of the sequence as is represented by the generated signal, and (ii) the same effective subsequences as are represented by the generated signal, or effective subsequences that are members of the same sets of effective subsequences as are represented by the generated signal,
whereby said one or more nucleic acids in said sample are identified, classified, or quantified.
119 . The method according to claim 110 which further comprises subjecting said pooled colonies of first nucleic acids to a method comprising a method for identifying, classifying, or quantifying DNA molecules in a sample, the method comprising the steps of:
(a) digesting said sample with one or more restriction endonucleases, each said restriction endonuclease recognizing a subsequence recognition site and digesting DNA at said recognition site to produce fragments with 5′ overhangs;
(b) contacting said produced fragments with shorter and longer oligodeoxynucleotides, each said shorter oligodeoxynucleotide hybridizable with a said 5′ overhang and having no terminal phosphates, each said longer oligodeoxynucleotide hybridizable with a said shorter oligodeoxynucleotide;
(c) ligating said longer oligodeoxynucleotides to said 5′ overhangs on said fragments to produce ligated DNA fragments;
(d) extending said ligated DNA fragments by synthesis with a DNA polymerase to produce blunt-ended double stranded DNA fragments;
(e) amplifying said blunt-ended double stranded DNA fragments by a method comprising contacting said blunt-ended double stranded DNA fragments with a DNA polymerase and primer oligodeoxynucleotides, each said primer oligodeoxyndcleotide having a sequence comprising that of one of the longer oligodeoxynucleotides;
(f) determining the length of the amplified DNA fragments produced in step (e); and
(g) searching a DNA sequence database, said database comprising a plurality of known DNA sequences that may be present in the sample, for sequences matching one or more of said fragments of determined length, a sequence from said database matching a fragment of determined length when the sequence from said database comprises recognition sites of said one or more restriction endonucleases spaced apart by the determined length,
whereby DNA molecules in said sample are identified, classified, or quantified.
120 . The method according to claim 111 which further comprises subjecting said pooled colonies of first DNA molecules to a method comprising a method for identifying, classifying, or quantifying DNA molecules in a sample, the method comprising the steps of:
(a) digesting said sample with one or more restriction endonucleases, each said restriction endonuclease recognizing a subsequence recognition site and digesting DNA at said recognition site to produce fragments with 5′ overhangs;
(b) contacting said produced fragments with shorter and longer oligodeoxynucleotides, each said shorter oligodeoxynucleotide hybridizable with a said 5′ overhang and having no terminal phosphates, each said longer oligodeoxynucleotide hybridizable with a said shorter oligodeoxynucleotide;
(c) ligating said longer oligodeoxynucleotides to said 5′ overhangs on said fragments to produce ligated DNA fragments;
(d) extending said ligated DNA fragments by synthesis with a DNA polymerase to produce blunt-ended double stranded DNA fragments;
(e) amplifying said blunt-ended double stranded DNA fragments by a method comprising contacting said blunt-ended double stranded DNA fragments with a DNA polymerase and primer oligodeoxynucleotides, each said primer oligodeoxynucleotide having a sequence comprising that of one of the longer oligodeoxynucleotides;
(f) determining the length of the amplified DNA fragments produced in step (e); and
(g) searching a DNA sequence database, said database comprising a plurality of known DNA sequences that may be present in the sample, for sequences matching one or more of said fragments of determined length, a sequence from said database matching a fragment of determined length when the sequence from said database comprises recognition sites of said one or more restriction endonucleases spaced apart by the determined length,
whereby DNA molecules in said sample are identified, classified, or quantified.
121 . The method according to claim 110 which further comprises subjecting said pooled colonies of second nucleic acids to a method comprising a method for identifying, classifying, or quantifying DNA molecules in a sample of DNA molecules having a plurality of different nucleotide sequences, the method comprising the steps of:
(a) digesting said sample with one or more restriction endonucleases, each said restriction endonuclease recognizing a subsequence recognition site and digesting DNA at said recognition site to produce fragments with 5′ overhangs;
(b) contacting said produced fragments with shorter and longer oligodeoxynucleotides, each said shorter oligodeoxynucleotide hybridizable with a said 5′ overhang and having no terminal phosphates, each said longer oligodeoxynucleotide hybridizable with a said shorter oligodeoxynucleotide;
(c) ligating said longer oligodeoxynucleotides to said 5′ overhangs on said fragments to produce ligated DNA fragments;
(d) extending said ligated DNA fragments by synthesis with a DNA polymerase to produce blunt-ended double stranded DNA fragments;
(e) amplifying said blunt-ended double stranded DNA fragments by a method comprising contacting said blunt-ended double stranded DNA fragments with a DNA polymerase and primer oligodeoxynucleotides, each said primer oligodeoxynucleotide having a sequence comprising that of one of the longer oligodeoxynucleotides;
(f) determining the length of the amplified DNA fragments produced in step (e); and
(g) searching a DNA sequence database, said database comprising a plurality of known DNA sequences that may be present in the sample, for sequences matching one or more of said fragments of determined length, a sequence from said database matching a fragment of determined length when the sequence from said database comprises recognition sites of said one or more restriction endonucleases spaced apart by the determined length,
whereby DNA molecules in said sample are identified, classified, or quantified.
122 . The method according to claim 111 which further comprises subjecting said pooled colonies of second DNA molecules to a method comprising a method for identifying, classifying, or quantifying DNA molecules in a sample of DNA molecules having a plurality of different nucleotide sequences, the method comprising the steps of:
(a) digesting said sample with one or more restriction endonucleases, each said restriction endonuclease recognizing a subsequence recognition site and digesting DNA at said recognition site to produce fragments with 5′ overhangs;
(b) contacting said produced fragments with shorter and longer oligodeoxynucleotides, each said shorter oligodeoxynucleotide hybridizable with a said 5′ overhang and having no terminal phosphates, each said longer oligodeoxynucleotide hybridizable with a said shorter oligodeoxynucleotide;
(c) ligating said longer oligodeoxynucleotides to said 5′ overhangs on said fragments to produce ligated DNA fragments;
(d) extending said ligated DNA fragments by synthesis with a DNA polymerase to produce blunt-ended double stranded DNA fragments;
(e) amplifying said blunt-ended double stranded DNA fragments by a method comprising contacting said blunt-ended double stranded DNA fragments with a DNA polymerase and primer oligodeoxynucleotides, each said primer oligodeoxynucleotide having a sequence comprising that of one of the longer oligodeoxynucleotides;
(f) determining the length of the amplified DNA fragments produced in step (e); and
(g) searching a DNA sequence database, said database comprising a plurality of known DNA sequences that may be present in the sample, for sequences matching one or more of said fragments of determined length, a sequence from said database matching a fragment of determined length when the sequence from said database comprises recognition sites of said one or more restriction endonucleases spaced apart by the determined length,
whereby DNA molecules in said sample are identified, classified, or quantified.
123 . The method according to claim 46 in which the population of diploid yeast cells is incubated in a third environment in which substantial death of yeast cells occurs in the absence of expression of the first and second selectable markers.
124 . Purified cells of a single yeast strain of mating type a, that is mutant in endogenous URA3 and HIS3, and contains functional URA3 coding sequences under the control of a promoter containing GAL4 binding sites, and contains functional lacZ coding sequences under the control of a promoter containing GAL4 binding sites.
125 . Purified cells of a single yeast strain of mating type α, that is mutant in endogenous URA3 and HIS3, and contains functional URA3 coding sequences under the control of a promoter containing GAL4 binding sites, and contains functional lacZ coding sequences under the control of a promoter containing GAL4 binding sites.
126 . A kit comprising in one or more containers:
(a) purified cells of a single yeast strain of mating type a, that is mutant in endogenous URA3 and HIS3, and contains functional URA3 coding sequences under the control of a promoter containing GAL4 binding sites, and contains functional lacZ coding sequences under the control of a promoter containing GAL4 binding sites; and (b) purified cells of a single yeast strain of mating type α, that is mutant in endogenous URA3 and HIS3, and contains functional URA3 coding sequences under the control of a promoter containing GAL4 binding sites, and contains functional lacZ coding sequences under the control of a promoter containing GAL4 binding sites.
127 . The kit of claim 126 which further comprises in one or more containers:
(c) a first vector comprising:
(i) a first promoter;
(ii) a first nucleotide sequence encoding a DNA binding domain, operably linked to the first promoter;
(iii) means for inserting a DNA sequence encoding a protein into the vector in such a manner that the protein is capable of being expressed as part of a fusion protein containing the DNA binding domain;
(iv) a transcription termination signal operably linked to the first nucleotide sequence;
(v) a first means for replicating in the cells of said yeast strains in (a) and (b); and
(d) a second vector comprising:
(i) a second promoter;
(ii) a nucleotide sequence encoding an activation domain of a transcriptional activator, operably linked to the second promoter;
(iii) means for inserting a DNA sequence encoding a protein into the vector in such a manner that the protein is capable of being expressed as part of a fusion protein containing the activation domain of a transcriptional activator;
(iv) a transcription termination signal operably linked to the second nucleotide sequence; and
(v) a second means for replicating in the purified cells of said yeast strains in (a) and (b).
128 . The method according to claim 50 in which said recombinant nucleic acids encoding said one or more candidate molecules each comprise the following operably linked components:
(a) an ADC1 promoter;
(b) a nucleotide sequence encoding a candidate molecule fused to a nuclear localization signal; and
(c) an ADC1 transcription termination signal.
129 . A purified expression vector comprising the following components:
(a) an ADC1 promoter; (b) a first nucleotide sequence encoding a nuclear localization signal, operably linked to the promoter; (c) means for inserting a DNA sequence into the vector in such a manner that a protein encoded by the DNA sequence is capable of being expressed as part of a fusion protein containing the nuclear localization signal; (d) an ADC1 transcription termination signal, operably linked to the first nucleotide sequence; (e) means for replicating in a yeast cell; (f) means for replicating in E. coli; (g) a second nucleotide sequence encoding a selectable marker for selection in a yeast cell, operably linked to a transcriptional promoter and transcription termination signal active in yeast; and (h) a third nucleotide sequence encoding a selectable marker for selection in E. coli , operably linked to a transcriptional promoter and transcription termination signal active in E. coli.
130 . A purified expression vector comprising the following components:
(a) a promoter active in yeast; (b) a first nucleotide sequence encoding a peptide of 20 or fewer amino acids fused to a nuclear localization signal, said first nucleotide sequence being operably linked to the promoter; (c) a transcription termination signal active in yeast, operably linked to said first nucleotide sequence; (d) means for replicating in a yeast cell; (e) means for replicating in E. coli; (f) a second nucleotide sequence encoding a selectable marker for selection in a yeast cell, operably linked to a transcriptional promoter and transcription termination signal active in yeast; and (g) a third nucleotide sequence encoding a selectable marker for selection in E. coli , operably linked to a transcriptional promoter and transcription termination signal active in E. coli.
131 . The method according to claim 50 in which diploid yeast cells have a mutation in at least one gene coding for a cell wall component thereby having a modified cell wall that is more permeable to exogenous molecules than is a wild-type cell wall.
132 . The method according to claim 54 in which said environment of incubating step (h) contains 3-amino-1,2,4-triazole.
133 . The method according to claim 123 in which said environment of incubating step (h) contains 3-amino-1,2,4-triazole.
134 . A method of detecting an inhibitor of a protein-protein interaction comprising
(a) incubating a population of cells, said population comprising cells recombinantly expressing a pair of interacting proteins, said pair consisting of a first fusion protein and a second fusion protein, in the presence of one or more candidate molecules among which it is desired to identify an inhibitor of the interaction between said first fusion protein and said second fusion protein, each said first fusion protein comprising a first protein sequence and a DNA binding domain; each said second fusion protein comprising a second protein sequence and a transcriptional activation domain of a transcriptional activator; and in which the cells contain a first nucleotide sequence operably linked to a promoter driven by one or more DNA binding sites recognized by said DNA binding domain such that an interaction of said first fusion protein with said second fusion protein results in increased transcription of said first nucleotide sequence, said incubating being in an environment in which substantial death of said cells occurs (i) when said increased transcription occurs of said first nucleotide sequence or (ii) if said cells lack a recombinant nucleic acid encoding said first fusion protein or a recombinant nucleic acid encoding said second fusion protein; and (b) detecting those cells that survive said incubating step, thereby detecting the presence of an inhibitor of said interaction in said cells.
135 . The method according to claim 134 in which said population of cells comprises a plurality of cells, each cell within said plurality recombinantly expressing a different said pair of interacting proteins.
136 . The method according to claim 134 in which the cells are yeast cells.
137 . The method according to claim 135 in which the cells are yeast cells.
138 . The method according to claim 136 or 137 in which the first nucleotide sequence is functional URA3 coding sequences, and said environment contains 5-fluoroorotic acid.
139 . The method according to claim 134 or 138 in which said one or more candidate molecules are provided to said cells by introducing into said cells one or more recombinant nucleic acids encoding said one or more candidate molecules, such that said one or more candidate molecules are expressed within said cells.
140 . The method according to claim 134 or 138 in which said environment contains said one or more candidate molecules.
141 . The method of claim 136 in which the cells are haploid yeast cells.
142 . The method of claim 136 in which the cells are diploid yeast cells.
143 . The method according to claim 17 in which cells are yeast cells, and in which the first protein and the second protein are first and second fusion proteins, respectively, between which an interaction is detected according to the method of claim 54 .
144 . The method according to claim 135 or 137 in which the plurality of cells consists of at least 10 cells.
145 . The method according to claim 135 or 137 in which the plurality of cells consists of at least 100 cells.
146 . The method according to claim 135 or 137 in which the plurality of cells consists of at least 1000 cells.
147 . The method according to claim 134 or 138 in which said one or more candidate molecules are compounds synthesized by proteins encoded by recombinant DNA that has been introduced into said cells.
148 . The method according to claim 17 in which said population of cells comprises a plurality of cells, each cell within said plurality recombinantly expressing a different said pair of interacting proteins.
149 . The method according to claim 148 in which the plurality of cells consists of at least 10 cells.
150 . The method according to claim 148 in which the plurality of cells consists of at least 100 cells.
151 . The method according to claim 148 in which the plurality of cells consists of at least 1000 cells.Join the waitlist — get patent alerts
Track US2003119002A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.