Nucleic acid-based data storage
Abstract
Methods and systems for encoding digital information in nucleic acid (e.g., deoxyribonucleic acid) molecules without base-by-base synthesis, by encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, comprising specifying each bit location in a bit-stream with a unique nucleic sequence and specifying the bit value at that location by the presence or absence of the corresponding unique nucleic acid sequence in the pool. But, more generally, specifying unique bytes in a bytestream by unique subsets of nucleic acid sequences. Also disclosed are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzymatic-based editing of nucleic acid sequences).
Claims
exact text as granted — not AI-modified1 - 29 . (canceled)
30 . A method for writing information into nucleic acid sequences, the method comprising:
translating the information into a string of symbols; mapping the string of symbols to a plurality of identifiers, wherein each individual identifier of the plurality of identifiers corresponds to a unique nucleic acid sequence and comprises a corresponding plurality of components, wherein each component of the plurality of components comprises a distinct nucleic acid sequence, and wherein each identifier corresponds to an individual symbol in the string of symbols; and forming at least one individual identifier of the plurality of identifiers by depositing the corresponding plurality of components into a compartment, wherein the plurality of components assemble via one or more reactions in the compartment to form the at least one individual identifier.
31 . The method of claim 30 , wherein each symbol is one of two possible symbol values.
32 . The method of claim 31 , wherein a first symbol value of the two possible symbol values is represented by an absence of a distinct identifier of the plurality of identifiers, and wherein a second symbol value of the two possible symbol values is represented by a presence of the distinct identifier, or vice versa.
33 . The method of claim 32 , wherein the two or more components assemble via the one or more reactions to form identifiers that each correspond to symbols having symbol values that are represented by the presence of the corresponding distinct identifiers.
34 . The method of claim 30 , wherein the distinct nucleic acid sequence of each individual component is non-specific to any individual symbol.
35 . The method of claim 30 , wherein each component comprises a distinct nucleic acid sequence with first and second ends, a first hybridization region on the first end, and a second hybridization region on the second end.
36 . The method of claim 30 , wherein each of the plurality of components belongs in one of M layers, and wherein one component from each of the M layers assemble to form the at least one identifier.
37 . The method of claim 36 , wherein each layer of the M layers comprises a distinct set of components.
38 . The method of claim 37 , wherein within each layer, the components have a common first hybridization region and a common second hybridization region.
39 . The method of claim 30 , wherein each component in the compartment has first and second hybridization regions, and the first or second hybridization region of each component is complementary to the first or second hybridization region of another component, and wherein the one or more reactions comprise hybridization of the first and second complimentary hybridization regions.
40 . The method of claim 30 , wherein the one or more components assemble in a non-order dependent manner.
41 . The method of claim 30 , wherein a subset of the plurality of identifiers are formed via one reaction in a multiplex fashion.
42 . The method of claim 30 , wherein the one or more reactions comprise overlap-extension polymerase chain reaction (PCR), polymerase cycling assembly, sticky end ligation, ligase cycling reaction, or template directed ligation.
43 . The method of claim 30 , further comprising generating an identifier library comprising the at least one formed identifier.
44 . The method of claim 43 , wherein the identifier library comprises a distinct barcode, metadata of the information, or both.
45 . The method of claim 44 , further comprising extracting a targeted subset of the identifier library.
46 . The method of claim 45 , further comprising combining a plurality of probes with the identifier library, wherein the plurality of probes share complementarity with the one or more components of each identifier of the targeted subset such that each identifier of the targeted subset hybridizes with at least one probe when combined with the plurality of probes.
47 . The method of claim 46 , wherein the plurality of probes comprises one or more affinity tags, and wherein the one or more affinity tags are captured by an affinity bead or an affinity column.
48 . The method of claim 45 , wherein each identifier comprises one or more common primer binding regions, one or more variable primer binding regions, or any combination thereof.
49 . The method of claim 48 , further comprising:
combining the identifier library with primers that bind to the one or more common primer binding regions or to the one or more variable primer binding regions, wherein the primers bind to the one or more common primer binding regions or to the one or more variable primer binding regions; and selectively amplifying said targeted subset of said identifier library.
50 . The method of claim 45 , further comprising selectively removing a portion of non-targeted identifiers from said identifier library.Join the waitlist — get patent alerts
Track US2020250546A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.