Error correction for nucleotide data stores
Abstract
This disclosure provides techniques for adding error correction to information in a data store that encodes information as a sequence of bases in polynucleotides. Errors may be introduced through creation of the database (e.g., oligonucleotide synthesis) and/or reading information from the database (e.g., polynucleotide sequencing). Additional polynucleotides added to the database can provide error correction through redundancy. The sequence of polynucleotides that provide error correction may be designed by performing an invertible summary operation on information to be stored in the database. One example of an invertible summary operation is the exclusive or operation (XOR). This disclosure also provides techniques for storing metadata related to organization of a database and structure of information on polynucleotides within the database. Metadata may be encoded in polynucleotides and added to the data store. The polynucleotides holding metadata may be designed with unique primer sites so that the metadata can be selectively amplified and sequenced.
Claims
exact text as granted — not AI-modified1 . A method of providing error correction for binary data encoded in synthetic polynucleotides, the method comprising:
synthesizing a first polynucleotide encoding a first information payload; synthesizing a second polynucleotide encoding a second information payload; and synthesizing a third polynucleotide encoding an error-correction payload that has less than full redundancy of the first information payload and less than full redundancy of the second information payload, wherein the error-correction payload is determined by an invertible summary operation on the first information payload and the second information payload.
2 . The method of claim 1 , wherein the first polynucleotide includes a first address encoded in a first identifier region of the first polynucleotide, the second polynucleotide includes a second address encoded a second identifier region of the second polynucleotide, and the third polynucleotide includes the first address and the second address encoded in a third identifier region of the third polynucleotide.
3 . The method of claim 2 , wherein the third identifier region of the third polynucleotide contains a nucleotide base sequence indicating that the third polynucleotide contains an error-correction payload.
4 . The method of claim 2 , further comprising synthesizing a fourth polynucleotide encoding metadata identifying the invertible summary operation.
5 . The method of claim 1 , wherein the invertible summary operation includes an exclusive or operation.
6 . The method of claim 1 , further comprising:
sequencing the first information payload, the second information payload, and the error-correction payload to identify a first nucleotide sequence of the first information payload, a second nucleotide sequence of the second information payload, and a third nucleotide sequence of the error-correction payload; converting the first nucleotide sequence, the second nucleotide sequence, and the third nucleotide sequence into a first binary data, a second binary data, and a third binary data; applying the invertible summary operation to the first binary data and the third binary data to generate a second instance of the second binary data; and correcting at least one error in the second binary data by comparing the second instance of the second binary data to the second binary data.
7 . A system for storing binary information in synthetic polynucleotides with error correction, the system comprising:
a first number of first polynucleotides, individual ones of the first polynucleotides encoding information payloads that represent binary data; and a second number of second polynucleotides, individuals ones of the second polynucleotides encoding error-correction payloads created by an invertible summary operation to have less than full redundancy of two or more of the information payloads.
8 . The system of claim 7 , wherein the first polynucleotides comprise one or more nucleotides identifying the first polynucleotides as encoding information payloads and the second polynucleotides comprise one or more nucleotides identifying the second polynucleotides as encoding error-correction payloads.
9 . The system of claim 8 , wherein the first polynucleotides include a first primer target, the second polynucleotides include a second primer target; and
further comprising a third polynucleotide having a third primer target different than the first primer target and different than the second primer target, the third polynucleotide encoding metadata describing the first primer target and the second primer target.
10 . The system of claim 7 , wherein the invertible summary operation includes an exclusive or operation.
11 . The system of claim 7 , wherein the second polynucleotides are physically separate from the first polynucleotides.
12 . The system of claim 7 , wherein a ratio of the first number to the second number is 2:1, 3:1, 4:1, 5:1, 3:2, or 5:2.
13 . The system of claim 7 , wherein the first number and the second number are selected based at least in part on a computer-readable file type associated with the binary information represented by the information payloads in the first polynucleotides.
14 . Computer storage media comprising instructions that when executed on a processor, cause the processor to perform acts comprising:
converting binary data into converted data represented as ternary data or quaternary data; encoding the converted data as a sequence of nucleotide bases; dividing the sequence of nucleotide bases into fragments, wherein a length a fragment is based at least in part on a length of polynucleotide that can be synthesized and an error rate associated with the length of polynucleotide; and generating an error-correction sequence of nucleotide bases by applying an invertible summary operation to at least two of the fragments, the error-correction sequence having less than full redundancy of the at least two of the fragments.
15 . The computer storage media of claim 14 , wherein the invertible summary operation is applied to one of the binary data, the converted data, or the sequence of nucleotide bases.
16 . The computer storage media of claim 14 , wherein the invertible summary operation includes an exclusive or operation.
17 . The computer storage media of claim 14 , wherein the acts further comprise creating a polynucleotide-synthesis template by appending a sequence of nucleotide bases representing a primer target and a sequence of nucleotide bases representing identifying information to a one of the fragments.
18 . The computer storage media of claim 17 , wherein the acts further comprise sending instructions to an oligonucleotide synthesizer to synthesize a polynucleotide having a nucleotide sequence represented by the polynucleotide-synthesis template.
19 . The computer storage media of claim 14 , wherein the acts further comprise:
encoding metadata as a sequence of nucleotide bases to create a metadata sequence; and creating a metadata polynucleotide-synthesis template by appending a metadata-specific primer target to the metadata sequence.
20 . The method of claim 19 , wherein the metadata comprises data describing one or more of a level of redundancy present for information in the data store, identity of the invertible summary operation, the technique used to convert binary data into a series of nucleic acid bases, polynucleotide length, payload region length, polynucleotide type, primer targets used in primary polynucleotides, or primer targets used in error-correction polynucleotides.Join the waitlist — get patent alerts
Track US2017141793A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.