Method for Achieving Improved Polypeptide Expression
Abstract
The present invention relates to methods of optimization of a protein coding sequences for expression in a given host cell. The methods apply genetic algorithms to optimise single codon fitness and/or codon pair fitness sequences coding for a predetermined amino acid sequence. In the algorithm generation of new sequence variants and subsequent selection of fitter variants is reiterated until the variant coding sequences reach a minimum value for single codon fitness and/or codon pair fitness. The invention also relates to a computer comprising a processor and memory, the processor being arranged to read from and write into the memory, the memory comprising data and instructions arranged to provide the processor with the capacity to perform the genetic algorithms for optimisation of single codon fitness and/or codon pair fitness. The invention further relates to nucleic acids comprising a coding sequence for a predetermined amino acid sequence, the coding sequence being optimised with respect to single codon fitness and/or codon pair fitness for a given host in the methods of the invention, to host cells comprising such nucleic acids and to methods for producing polypeptides and other fermentation products in which these host cells are used.
Claims
exact text as granted — not AI-modified1 . An isolated nucleic acid molecule comprising a coding sequence coding for a predetermined amino acid sequence, wherein the coding sequence is not a naturally occurring coding sequence and wherein the coding sequence has a fit cp (g) of below −0.1 for a predetermined host cell.
2 . An isolated nucleic acid molecule according to claim 1 , having a fit sc (g) of below 0.1 for the predetermined host cell.
3 . An isolated nucleic acid molecule according to claim 1 , having a fit cp (g) of below −0.2 for the predetermined host cell.
4 . An isolated nucleic acid molecule according to claim 3 , a fit sc (g) of below 0.1 for a predetermined host cell.
5 . The isolated nucleic acid of claim 1 , wherein said coding sequence is generated by a method comprising:
(a) generating at least one coding sequence that codes for the predetermined amino acid sequence; (b) generating at least one newly generated coding sequence from at least one coding sequence by replacing in at least one coding sequence at least one codons by a synonymous codon; (c) determining a fitness value of said at least one coding sequence and a fitness value of said at least one newly generated coding sequence while using a fitness function that determines at least one of single codon fitness and codon pair fitness for the predetermined host cell; (d) choosing at least one selected coding sequence amongst said at least one coding sequence and said at least one newly generated coding sequence in accordance with a predetermined selection criterion such that the higher is said fitness value, the higher is a chance of being chosen; (e) repeating (b) through (d) while treating said at least one selected coding sequence as at least one coding sequence in (b) through (d) until a predetermined iteration stop criterion is fulfilled.
6 . The isolated nucleic acid according to claim 5 , wherein said predetermined selection criterion is such that said at least one selected coding sequence have a best fitness value according to a predetermined criterion.
7 . The isolated nucleic acid according to claim 5 , wherein said method of generating the coding sequence comprises, after (e):
(f) selecting a best individual coding sequence amongst said at least one selected coding sequences where said best individual coding sequence has a better fitness value than other selected coding sequences.
8 . The isolated nucleic acid according to claim 5 , wherein said predetermined iteration stop criterion is at least one of the following:
(e1) testing whether at least one of said selected coding sequences have a best fitness value above a predetermined threshold value; or (e2) testing whether none of said selected coding sequences has a best fitness value below said predetermined threshold value; or (e3) testing whether at least one of said selected coding sequences has at least 30% of the codon pairs with associated positive codon pair weights for the predetermined host cell in said coding sequence being transformed into codon pairs with associated negative weights; or (e4) testing whether at least one of said selected coding sequences has at least 30% of the codon pairs with associated positive weights above 0 for the predetermined host cell in said coding sequence being transformed into codon pairs with associated weights below 0.
9 . The isolated nucleic acid according to claim 5 , wherein said coding nucleotide sequence coding for a predetermined amino acid sequence is selected from the group consisting of:
(a1) a wild-type nucleotide sequence coding for said predetermined amino acid sequence; (a2) a reverse translation of the predetermined amino acid sequence whereby a codon for an amino acid position in the predetermined amino acid sequence is randomly chosen from the synonymous codons coding for the amino acid; and (a3) a reverse translation of the predetermined amino acid sequence whereby a codon for an amino acid position in the predetermined amino acid sequence is chosen in accordance with a single-codon bias for the predetermined host cell or a species related to the host cell.
10 . The isolated nucleic acid of claim 5 , wherein said fitness function defines single codon fitness by:
fit
c
(
g
)
=
100
-
1
g
·
∑
k
=
1
g
r
c
target
(
c
(
k
)
)
-
r
c
g
(
c
(
k
)
)
·
100
wherein g symbolizes a coding sequence, |g| is the length of the coding sequence g, g(k) is the k-th codon of the coding sequence g, r c target (c(k) is a desired ratio of codon c(k), and r c g (c(k)) is an actual ratio in the nucleotide coding sequence g.
11 . The isolated nucleic acid of claim 5 , wherein said fitness function defines codon pair fitness:
fit
cp
(
g
)
=
1
g
-
1
·
∑
k
=
1
g
-
1
w
(
(
c
(
k
)
,
c
(
k
+
1
)
)
wherein w((c(k), c(k+1)) is a weight of a codon pair in a coding sequence g, |g| is length of said nucleotide coding sequence and c(k) is k-th codon in said coding sequence.
12 . The isolated nucleic acid according to claim 11 , wherein said codon pair weights w are taken from a 61×61 codon pair matrix without stop codons, or a 61×64 codon pair matrix that includes stop-codons, and wherein said codon pair weights w are calculated using at least one of the following:
(a) a group of nucleotide sequences consisting of at least 200 coding sequences of a predetermined host; or
(b) a group of nucleotide sequences consisting of at least 200 coding sequences of the species to which the predetermined host belongs; or
(c) a group of nucleotide sequences consisting of at least 5% of the protein encoding nucleotide sequences in a genome sequence of the predetermined host; or
(d) a group of nucleotide sequences consisting of at least 5% of the protein encoding nucleotide sequences in a genome sequence of a genus related to the predetermined host.
13 . A method according to claim 12 , wherein said codon pair weights w are determined for at least 50% of the possible 61×64 codon pairs including the termination signal as stop codon
14 . The isolated nucleic acid of claim 11 , wherein said fitness function is defined by:
fit
combi
(
g
)
=
fit
cp
(
g
)
cpi
+
fit
sc
(
g
)
wherein
fit
sc
(
g
)
=
1
g
·
∑
k
=
1
g
r
sc
target
(
c
(
k
)
)
-
r
sc
g
(
c
(
k
)
)
cpi is a real value greater than zero, fit sc (g) is a single codon fitness function, r sc target (c(k)) is a desired ratio of codon c(k), and r sc g (c(k)) is an actual ratio in the coding sequence g.
15 . The isolated nucleic acid according to claim 14 , wherein cpi is between 10 −4 and 0.5.
16 . The isolated nucleic acid according to claim 1 , wherein said predetermined host cell is a cell of a microorganism of a genus selected from the group consisting of: Bacillus, Actinomycetis, Escherichia, Streptomyces, Aspergillus, Penicillium, Kluyveromyces , and Saccharomyces.
17 . The isolated nucleic acid according to claim 1 , wherein said predetermined host cell is a cell of an animal or plant.
18 . The isolated host cell of claim 17 , wherein said predetermined host cell is a cell of a cell line selected from the group consisting of CHO, BHK, NS0, COS, Vero, PER.C6™, HEK-293, Drosophila S2, Spodoptera Sf9, and Spodoptera Sf21.
19 . The isolated nucleic acid molecule according to claim 1 , wherein the coding sequence is operably linked to at least one expression control sequence capable of directing expression of the coding sequence in the predetermined host cell.
20 . A host cell comprising an isolated nucleic acid molecule as defined in claim 19 .
21 . A method for producing a polypeptide having the predetermined amino acid sequence, the method comprising culturing a host cell as defined in claim 20 under conditions conducive to the expression of the polypeptide.
22 . The method of claim 21 further comprising recovering the polypeptide.
23 . A method for producing at least one of an intracellular and an extracellular metabolite, the method comprising culturing a host cell as defined in claim 20 under conditions conducive to the production of the metabolite.
24 . The method of claim 23 , wherein the polypeptide having the predetermined amino acid sequence is involved in the production of the metabolite.
25 . The isolated nucleic acid according to claim 11 , wherein said codon pair weights w are taken from a 61×61 codon pair matrix without stop codons, or a 61×64 codon pair matrix that includes stop-codons, and wherein said codon pair weights w are defined by means of:
w
(
(
c
i
,
c
j
)
)
=
n
exp
combi
(
(
c
i
,
c
j
)
)
-
n
obs
high
(
(
c
i
,
c
j
)
)
max
(
n
obs
high
(
(
c
i
,
c
j
)
)
,
n
exp
combi
(
(
c
i
,
c
j
)
)
)
where the combined expected values n exp combi ((c i ,c j )) defined by means of:
n
exp
combi
(
(
c
i
,
c
j
)
)
=
r
sc
all
(
c
i
)
·
r
sc
all
(
c
j
)
·
∑
c
k
∈
syn
(
c
i
)
c
l
∈
syn
(
c
j
)
n
obs
high
(
(
c
k
,
c
l
)
)
where r sc all (c k ) denote the single codon ratio of c k in the whole genome data set and n obs high (c i ,c j )) the occurrences of a pair (c i ,c j ) in the highly expressed group, and wherein the highly expressed group are the genes whose mRNA's can be detected at a level of at least 20 copies per cell.Join the waitlist — get patent alerts
Track US2014377800A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.