US2025391507A1PendingUtilityA1

Embedded Reference Marks for Correcting Errors in DNA Data Storage

Assignee: WESTERN DIGITAL TECH INCPriority: Jun 21, 2024Filed: Jun 21, 2024Published: Dec 25, 2025
Est. expiryJun 21, 2044(~17.9 yrs left)· nominal 20-yr term from priority
G06F 17/16G16B 50/50G16B 30/10
56
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Example systems and methods for using embedded reference marks and a correlation matrix to correct insertions and deletions for DNA data storage are described. A data unit may be encoded in oligos that include reference marks at predetermined intervals along the length of each oligo. During decoding, a comparison of reference marks from the read data of the oligo to a known reference mark pattern may be used to populate a correlation matrix. A most likely path for traversing the correlation matrix may be determined to identify offsets corresponding to insertions and deletions in the oligo, which may then be corrected during further decoding of the oligo.

Claims

exact text as granted — not AI-modified
1 . A system, comprising: 
 a decoder configured to: 
 receive read data determined from sequencing of an oligo for encoding a data unit, wherein the oligo comprises: 
 a number of symbols corresponding to user data in the data unit; and 
 a plurality of reference marks encoded at a predetermined interval along a length of the oligo; 
 
 populate a correlation matrix based on a comparison of a known reference mark pattern to reference mark positions in the read data; 
 determine a most likely path through the correlation matrix corresponding to offset values for the plurality of reference marks; 
 determine, based on at least one offset value from the most likely path, an error in a data segment between sequential reference marks; 
 correct symbol alignment in the data segment to compensate for error;  
 decode the user data from the read data; and 
 output, based on the decoded user data, the data unit. 
   
     
     
         2 . The system of  claim 1 , wherein the plurality of reference marks comprises a predetermined sequence of base pairs corresponding to the known reference mark pattern inserted at the predetermined intervals along the length of the oligo during encoding. 
     
     
         3 . The system of  claim 1 , wherein determining the most likely path through the correlation matrix comprises applying a Viterbi algorithm to traverse the correlation matrix. 
     
     
         4 . The system of  claim 1 , wherein the predetermined interval of the plurality of reference marks corresponds to a single base pair between sequential reference marks. 
     
     
         5 . The system of  claim 1 , wherein the comparison of reference mark positions is based on comparing: 
 a convolutional matrix comprised of the read data, wherein each column of the convolutional matrix corresponds to a base pair offset of the read data; and   a reference matrix comprised of the known reference mark pattern, wherein each column of the reference matrix repeats values from reference mark positions of the known reference mark pattern.   
     
     
         6 . The system of  claim 1 , wherein the correlation matrix comprises: 
 rows corresponding to a sequence of a subset of positions along the oligo corresponding to encoded reference mark positions at the predetermined interval;   columns corresponding to single base pair shifts in relative positions of the read data and the known reference mark pattern; and   matrix values corresponding to an exclusive-or comparison of corresponding base pairs of the read data and the known reference mark pattern.   
     
     
         7 . The system of  claim 1 , wherein determining the most likely path through the correlation matrix comprises: 
  traversing the correlation matrix to determine a series of probabilities of changing from a current column to an adjacent column for each row; and    calculating, based on the series of probabilities, a path having a highest likelihood among possible paths.   
     
     
         8 . The system of  claim 7 , wherein: 
 traversing the correlation matrix comprises: 
 traversing the correlation matrix in a first direction across the correlation matrix to determine forward probabilities; and 
 traversing the correlation matrix in an opposite direction across the correlation matrix to determine reverse probabilities; 
   calculating the path having the highest likelihood among possible paths uses a summation of the forward probabilities and the reverse probabilities; and   determining the forward probabilities and the reverse probabilities is based on a Toeplitz matrix.    
     
     
         9 . The system of  claim 7 , wherein: 
 the decoder is further configured to determine, based on the series of probabilities, soft information for the plurality of reference marks and adjacent user data positions between sequential reference marks; and   decoding the user data from the read data comprises using a user data decoder configured to: 
 receive the soft information from determining the most likely path through the correlation matrix; and 
 decode, using the soft information, the number of symbols corresponding to the user data in the read data. 
   
     
     
         10 . The system of  claim 1 , further comprising: 
 an encoder configured to: 
 determine the oligo for encoding the data unit; 
 determine the plurality of reference marks corresponding to the known reference mark pattern; 
 insert the plurality of reference marks at the predetermined interval along the length of the oligo, wherein the predetermined interval of the plurality of reference marks corresponds to a single base pair between sequential reference marks; and 
 output write data for the oligo for synthesis of the oligo. 
   
     
     
         11 . A method comprising: 
 receiving read data determined from sequencing an oligo that encodes a data unit, wherein the oligo comprises: 
 a number of symbols corresponding to user data in the data unit; and 
 a plurality of reference marks encoded at a predetermined interval along a length of the oligo; 
   populating a correlation matrix based on a comparison of reference mark positions in the read data to a known reference mark pattern;   determining a most likely path through the correlation matrix corresponding to offset values for the plurality of reference marks;   determining, based on at least one offset value from the most likely path, an error in a data segment between sequential reference marks;   correcting symbol alignment in the data segment to compensate for the insertion or deletion;    decoding the user data from the read data; and   outputting, based on the decoded user data, the data unit.   
     
     
         12 . The method of  claim 11 , wherein the plurality of reference marks comprises a predetermined sequence of base pairs corresponding to the known reference mark pattern inserted at the predetermined intervals along the length of the oligo during encoding. 
     
     
         13 . The method of  claim 11 , wherein determining the most likely path through the correlation matrix comprises applying a Viterbi algorithm to traverse the correlation matrix. 
     
     
         14 . The method of  claim 11 , wherein the predetermined interval of the plurality of reference marks corresponds to a single base pair. 
     
     
         15 . The method of  claim 11 , wherein the comparison of reference mark positions is based on: 
 a convolutional matrix comprised of the read data, wherein each column of the convolutional matrix corresponds to a base pair offset of the read data; and   a reference matrix comprised of the known reference mark pattern, wherein each column of the reference matrix repeats values from reference mark positions of the known reference mark pattern.   
     
     
         16 . The method of  claim 11 , wherein the correlation matrix comprises: 
 rows corresponding to a sequence of a subset of positions along the oligo corresponding to encoded reference mark positions at the predetermined interval;   columns corresponding to single base pair shifts in relative positions of the read data and the known reference mark pattern; and   matrix values corresponding to an exclusive-or comparison of corresponding base pairs of the read data and the known reference mark pattern.   
     
     
         17 . The method of  claim 11 , wherein determining the most likely path through the correlation matrix comprises: 
 traversing the correlation matrix to determine a series of probabilities of changing from a current column to an adjacent column for each row; and   calculating, based on the series of probabilities, a path having a highest likelihood among possible paths.   
     
     
         18 . The method of  claim 17 , wherein: 
 traversing the correlation matrix comprises: 
 traversing the correlation matrix in a first direction across the correlation matrix to determine forward probabilities; and 
 traversing the correlation matrix in an opposite direction across the correlation matrix to determine reverse probabilities; 
   calculating the path having the highest likelihood among possible paths uses a summation of the forward probabilities and the reverse probabilities; and   determining the forward probabilities and the reverse probabilities is based on a Toeplitz matrix.    
     
     
         19 . The method of  claim 17 , further comprising: 
 determining, based on the series of probabilities, soft information for the plurality of reference marks and adjacent user data positions between sequential reference marks, wherein decoding the user data from the read data comprises: 
 receiving, by a user data decoder, the soft information from determining the most likely path through the correlation matrix; and 
 decoding, using the soft information, the number of symbols corresponding to the user data in the read data. 
   
     
     
         20 . A system comprising: 
 means for receiving read data determined from sequencing an oligo that encodes a data unit, wherein the oligo comprises: 
 a number of symbols corresponding to user data in the data unit; and 
 a plurality of reference marks encoded at a predetermined interval along a length of the oligo; 
   means for populating a correlation matrix based on a comparison of reference mark positions in the read data to a known reference mark pattern;   means for determining a most likely path through the correlation matrix corresponding to offset values for the plurality of reference marks;   means for determining, based on at least one offset value from the most likely path, an error in a data segment between sequential reference marks;   means for correcting symbol alignment in the data segment to compensate for the insertion or deletion;    means for decoding the user data from the read data; and   means for outputting, based on the decoded user data, the data unit.

Join the waitlist — get patent alerts

Track US2025391507A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.