US2011269119A1PendingUtilityA1

Encoding text into nucleic acid sequences

Assignee: SYNTHETIC GENOMICS INCPriority: Oct 30, 2009Filed: Oct 29, 2010Published: Nov 3, 2011
Est. expiryOct 30, 2029(~3.2 yrs left)· nominal 20-yr term from priority
C12Q 1/6806C12Q 1/68G16B 30/00G06F 16/86
50
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods and apparatus are disclosed herein for encoding human readable text conveying a non-genetic message into nucleic acid sequences with a substantially reduced probability of biological impact and decoding such text from nucleic acid sequences. In one embodiment, each symbol of a symbol set of human readable symbols uniquely maps to a respective codon identifier. Mapping may ensure that each symbol will not map to a codon identifier that generates an amino acid residue which has a single-letter abbreviation that is the equivalent to the respective symbol. Synthetic nucleic acid sequences comprising such human readable text, and recombinant or synthetic cells comprising such sequences are provided, as well as methods of identifying cells, organisms, or samples containing such sequences.

Claims

exact text as granted — not AI-modified
1 . A synthetic nucleic acid sequence, wherein said synthetic nucleic acid sequence comprises one or more codon identifiers corresponding to a set of human readable symbols of a reference language, and further wherein said synthetic nucleic acid sequence is not genetically viable; does not have a biological impact upon a recombinant or synthetic organism comprising said synthetic nucleic acid sequence; and conveys a non-genetic message. 
     
     
         2 . The synthetic nucleic acid sequence of  claim 1 , wherein said synthetic nucleic acid sequence cannot be biologically translated into a functional amino acid sequence by the recombinant or synthetic organism. 
     
     
         3 . The synthetic nucleic acid sequence of  claim 1 , wherein said one or more codon identifiers:
 (i) do not correspond to sequence of a gene or other biologically active sequence;   (ii) correspond to one or more letters, one or more numbers, one or more spaces, one or more punctuation marks, one or more mathematical symbols, one or more typographical characters, one or more new lines, or a combination of any thereof; or   (iii) consists of three nucleotides.   
     
     
         4 . The synthetic nucleic acid sequence of  claim 1 , wherein said set of human readable symbols comprises a watermark that is a copyright notice, a trademark, a company identifier, a name, a phrase, a sentence, a quotation, genetic information, unique identifying information, data, or a combination of any thereof. 
     
     
         5 . The synthetic nucleic acid sequence of  claim 1 , further comprising an all-6 reading frame stop codon containing sequence 5′ to a first codon identifier in the sequence and/or an all-6 reading frame stop codon containing sequence 3′ to the last codon identifier in the sequence. 
     
     
         6 . A recombinant or synthetic organism comprising the synthetic nucleic acid sequence of  claim 1 . 
     
     
         7 . The recombinant or synthetic organism of  claim 9 , wherein said recombinant or synthetic organism is a prokaryotic cell, a eukaryotic cell, an archael cell or a virus. 
     
     
         8 . A method of creating a recombinant or synthetic organism comprising a watermark that conveys a non-genetic message, said method comprising:
 (i) generating a nucleic acid sequence comprising a sequence of codon identifiers selected based upon the text of the watermark such that a symbol mapping maps codon identifiers corresponding to start codon(s) to human readable symbols that possess a disproportionally low frequency in the language of the watermark, and maps codon identifiers corresponding to stop codon(s) to human readable symbols that possess a disproportionally high frequency in the language of the watermark;   (ii) synthesizing said nucleic acid sequence; and   (iii) introducing said nucleic acid sequence into a recombinant or synthetic organism, thereby creating said recombinant or synthetic organism comprising a watermark.   
     
     
         9 . The method of  claim 8 , wherein the symbol mapping does not map a three nucleotide codon identifier to a single letter representation of an amino acid residue normally assigned to that three nucleotide codon in the standard genetic code. 
     
     
         10 . The method of  claim 8 , wherein said generating step (i) is computer-assisted and comprises identifying the set of human readable symbols at a memory module and for each human readable symbol in the set, and using a processor to read a symbol mapping for determining a codon identifier which maps to the respective human readable symbol. 
     
     
         11 . A method of determining the presence of one or more recombinant or synthetic organisms comprising a reference watermark that conveys a non-genetic message in a sample, said method comprising:
 (i) sequencing nucleic acid material obtained from one or more organisms in said sample;   (ii) transforming the nucleic acid sequence obtained in step (i) to a set of codon identifiers, wherein each codon identifier of said set of codon identifiers consists of three nucleotides, and said transforming is performed in all three reading frames;   (iii) determining a human readable symbol for each codon identifier in the sequence in all three reading frames, wherein said determination is based at least in part upon a symbol mapping that map codons identifiers corresponding to start codon(s) to human readable symbols that possess a disproportionally low frequency in the language of the watermark, and that maps codon identifiers corresponding to stop codon(s) to human readable symbols that possess a disproportionally high frequency in the language of the watermark; and   (iv) comparing the human readable symbol sequence of all three reading frames to the reference watermark in said recombinant or synthetic organism, whereby the presence of the reference watermark in any reading frame of the nucleic acid material obtained in step (i) indicates the presence of the one or more recombinant or synthetic organism in the sample.   
     
     
         12 . An apparatus for transforming a sequence of codon identifiers into a sequence of human readable symbols that conveys a non-genetic message, the apparatus comprising:
 (i) a processor adapted to execute instructions; and   (ii) a storage module, wherein the storage module comprises a data structure for mapping codon identifiers into human readable symbols, and a set of instructions which, when executed by the processor, generate a human readable symbol for each codon identifier read from a sequence of codon identifiers, wherein the human readable symbol generated is based at least in part upon the data structure;   wherein the data structure is configured to map a start codon to a human readable symbol with a frequency of occurrence within a reference language that is less than a first predetermined threshold, and wherein the data structure is further configured to map a plurality of stop codons to human readable symbols with frequencies of occurrence within the reference language that are greater than a second predetermined threshold.   
     
     
         13 . The apparatus of  claim 12 , wherein the data structure does not map a codon identifier to a single letter representation of an amino acid residue normally assigned to that codon identifier in the standard genetic code. 
     
     
         14 . The apparatus of  claim 12 , wherein the sequence of codon identifiers comprises an all-6 reading frame stop codon containing sequence 5′ to a first codon identifier in the sequence and/or an all-6 reading frame stop codon containing sequence 3′ to the last codon identifier in the sequence. 
     
     
         15 . A method of transforming a first signal adapted to indicate a sequence of codon identifiers into a second signal adapted to indicate a sequence of human readable symbols that conveys a non-genetic message, the method comprising:
 (i) receiving the first signal;   (ii) determining a human readable symbol for each codon identifier in the sequence, wherein said determining is performed in all three reading frames and is based at least in part upon a mapping function configured to map a start codon to a first human readable symbol, wherein the first human readable symbol has a lower frequency of occurrence in a human readable symbol sequence than one or more symbols from a set of human readable symbols containing the first human readable symbol, and wherein the mapping function is further configured to map a stop codon to a second human readable symbol, wherein the second human readable symbol is contained within the set of human readable symbols, and wherein the second human readable symbol has a higher frequency of occurrence in the human readable symbol sequence than one or more human readable symbols from the set of human readable symbols; and   (iii) transforming the first signal into the second signal based upon the one or more determined human readable symbols.   
     
     
         16 . The method of  claim 15 , wherein the mapping function does not map a codon identifier to a single letter representation of an amino acid residue normally assigned to that codon identifier in the standard genetic code. 
     
     
         17 . The method of  claim 15 , wherein the sequence of codon identifiers comprises an all-6 reading frame stop codon containing sequence 5′ to a first codon identifier in the sequence and/or an all-6 reading frame stop codon containing sequence 3′ to the last codon identifier in the sequence. 
     
     
         18 . An apparatus for converting a sequence of human readable symbols of a reference language that conveys a non-genetic message into a sequence of codon identifiers, the apparatus comprising:
 (i) a processor configured to execute instructions;   (ii) a memory module coupled to the processor and comprising instructions which, when executed by the processor, determine a codon identifier for each human readable symbol contained within the sequence of human readable symbols, wherein each codon identifier is determined upon reading a symbol map; and   (iii) a data module coupled to the memory module,   wherein the data module comprises the symbol map, wherein the symbol map is configured to map one or more start codons to respective human readable symbols that possess a disproportionally low frequency of occurrence in the reference language, and wherein the symbol map is further configured to map one or more stop codons to respective human readable symbols that possess a disproportionally high frequency in the reference language.   
     
     
         19 . A computer-readable medium for use in an encoding machine, the computer-readable medium comprising instructions which, when executed by the encoding machine, perform a process comprising:
 (i) receiving a sequence of human readable symbols that conveys a non-genetic message; and   (ii) generating a codon identifier for each human readable symbol contained within the sequence,   wherein the human readable symbol generated is based at least in part upon a mapping function configured to map a start codon to a first human readable symbol, wherein the first human readable symbol has a lower frequency of occurrence in a reference language than one or more human readable symbols from a set of human readable symbols containing the first human readable symbol, and wherein the mapping function is further configured to map a stop codon to a second human readable symbol, wherein the second human readable symbol is contained within the set of human readable symbols, and wherein the second human readable symbol has a higher frequency of occurrence in the reference language than one or more human readable symbols from the set of human readable symbols.   
     
     
         20 . A method of generating a sequence of codon identifiers from a sequence of human readable symbols that conveys a non-genetic message, the method comprising:
 (i) receiving the sequence of human readable symbols at a memory module;   (ii) loading a human readable symbol map within the memory module, wherein the human readable symbol map is configured to determine a codon identifier that maps to each human readable symbol within the sequence, wherein the human readable symbol map is further configured to map a human readable symbol with a frequency of occurrence that is less than a first predetermined threshold within a reference language to a start codon, and wherein the symbol map is further configured to map a human readable symbol with a frequency of occurrence that is greater than a second predetermined threshold within the reference language to a stop codon; and   (iii) outputting a sequence of codon identifiers corresponding to each human readable symbol within the sequence.

Join the waitlist — get patent alerts

Track US2011269119A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.