US12183435B2ActiveUtilityA1

Systems and methods using DNA sequence strings as a common data format for forensic DNA typing applications

Assignee: BATTELLE MEMORIAL INSTITUTEPriority: Jul 30, 2016Filed: Apr 17, 2023Granted: Dec 31, 2024
Est. expiryJul 30, 2036(~10 yrs left)· nominal 20-yr term from priority
Inventors:Mark R. Wilson
G16B 50/00G06F 16/13G06F 16/148G16B 30/00
83
PatentIndex Score
0
Cited by
9
References
20
Claims

Abstract

The present disclosure relates, generally, to nucleotide sequence data and, more particularly, to computer files and methods supporting forensic DNA analysis. In one illustrative embodiment, a method may comprise identifying a locus corresponding to each item of short tandem repeat (STR) profiling data stored in an existing computer file, wherein the STR profiling data stored in the existing computer file is repeat-based and/or length-based; identifying start and stop coordinates of an STR region of the corresponding locus for each item of STR profiling data stored in the existing computer file; creating an ambiguous text string corresponding to each item of STR profiling data stored in the existing computer file, wherein each ambiguous text string consists of a sequence of ambiguous characters extending from the start coordinate to the stop coordinate identified for the corresponding item of STR profiling data; and storing each ambiguous text string in a sequence-based computer file.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A method comprising:
 creating an ambiguous text string corresponding to each item of short tandem repeat (STR) profiling data stored in an existing computer file, wherein each ambiguous text string consists of a sequence of ambiguous characters extending from a start coordinate to a stop coordinate identified for the corresponding item of STR profiling data, wherein each ambiguous character of each ambiguous text string is a text character used to represent an ambiguous nucleotide base call; and 
 storing each ambiguous text string in a sequence-based computer file. 
 
     
     
       2. The method of  claim 1 , wherein the STR profiling data stored in the existing computer file is repeat-based and/or length-based. 
     
     
       3. The method of  claim 1 , further comprising obtaining the existing computer file storing the short tandem repeat (STR) profiling data from one of the Combined DNA Index System (CODIS) database or the United Kingdom National DNA Database (NDNAD). 
     
     
       4. The method of  claim 1 , further comprising storing, in the sequence-based computer file, data representing the start and stop coordinates identified for each item of STR profiling data together with the ambiguous text string created for each item of STR profiling data. 
     
     
       5. The method of  claim 1 , wherein the start and stop coordinates identified for each item of STR profiling data are the start and stop coordinates of an STR region of a corresponding locus identified for that item of STR profiling data. 
     
     
       6. The method of  claim 5 , further comprising storing, in the sequence-based computer file, data representing the corresponding locus identified for each item of STR profiling data together with the ambiguous text string created for each item of STR profiling data. 
     
     
       7. The method of  claim 1 , further comprising:
 repeating, for a plurality of existing computer files storing repeat-based and/or length-based STR profiling data, the steps of (i) creating an ambiguous text string corresponding to each item of STR profiling data stored in the existing computer file and (ii) storing each ambiguous text string in a sequence-based computer file, and 
 storing the resulting plurality of sequence-based computer files in a database. 
 
     
     
       8. A non-transitory computer readable medium storing sequence-based short tandem repeat (STR) profiling data for a DNA sample, the sequence-based STR profiling data comprising a plurality of ambiguous text strings each representing an STR region of a locus of the DNA sample, wherein each ambiguous text string consists of a sequence of ambiguous characters extending from a start coordinate to a stop coordinate of the STR region represented by that ambiguous text string, wherein each ambiguous character of each ambiguous text string is a text character used to represent an ambiguous base call. 
     
     
       9. The non-transitory computer readable medium of  claim 8 , wherein the sequence-based STR profiling data further comprises locus data associated with each ambiguous text string of the plurality of ambiguous text strings, wherein the locus data identifies the locus containing the STR region represented by the associated ambiguous text string. 
     
     
       10. The non-transitory computer readable medium of  claim 8 , wherein the sequence-based STR profiling data further comprises coordinate data associated with each ambiguous text string of the plurality of ambiguous text strings, wherein the coordinate data identifies the start and stop coordinates of the STR region represented by the associated ambiguous text string. 
     
     
       11. The non-transitory computer readable medium of  claim 8 , wherein the DNA sample belongs to a human. 
     
     
       12. A non-transitory computer readable medium storing a plurality of instructions configured to cause a processor that executes the plurality of instructions to:
 receive a target text string that represents nucleotide sequence data generated by a read of a massively parallel sequencing (MPS) instrument; and 
 compare the target text string to each of the plurality of ambiguous text strings of the sequence-based STR profiling data for the DNA sample stored on the non-transitory computer readable medium of  claim 8 . 
 
     
     
       13. A non-transitory computer readable medium storing a plurality of instructions configured to cause a processor that executes the plurality of instructions to:
 receive a plurality of target text strings, wherein each of the plurality of target text strings represents nucleotide sequence data generated by a read of a massively parallel sequencing (MPS) instrument; and 
 compare the plurality of target text strings to the plurality of ambiguous text strings of the sequence-based STR profiling data for the DNA sample stored on the non-transitory computer readable medium of  claim 8 . 
 
     
     
       14. The non-transitory computer readable medium of  claim 13 , wherein the plurality of instructions is configured to cause the processor to compare the plurality of target text strings to the plurality of ambiguous text strings by comparing each target text string to each ambiguous text string that is associated with the same locus as that target text string. 
     
     
       15. A non-transitory computer readable medium storing a database of sequence-based short tandem repeat (STR) profiling data for a plurality of DNA samples, each of the plurality of DNA samples having a corresponding database record comprising a plurality of ambiguous text strings that each represent an STR region of a locus of the corresponding DNA sample, wherein each ambiguous text string consists of a sequence of ambiguous characters extending from a start coordinate to a stop coordinate of the STR region represented by that ambiguous text string, wherein each ambiguous character of each ambiguous text string is a text character used to represent an ambiguous base call. 
     
     
       16. The non-transitory computer readable medium of  claim 15 , wherein each of the plurality of DNA samples belongs to a different human. 
     
     
       17. A non-transitory computer readable medium storing a plurality of instructions configured to cause a processor that executes the plurality of instructions to:
 receive a target text string that represents nucleotide sequence data generated by a read of a massively parallel sequencing (MPS) instrument; and 
 compare the target text string to each of the plurality of ambiguous text strings in each database record of the database stored on the non-transitory computer readable medium of  claim 15 . 
 
     
     
       18. A non-transitory computer readable medium storing a plurality of instructions configured to cause a processor that executes the plurality of instructions to:
 receive a plurality of target text strings, wherein each of the plurality of target text strings represents nucleotide sequence data generated by a read of a massively parallel sequencing (MPS) instrument; and 
 compare the plurality of target text strings to the plurality of ambiguous text strings in each database record of the database stored on the non-transitory computer readable medium of  claim 15 . 
 
     
     
       19. The non-transitory computer readable medium of  claim 18 , wherein the plurality of instructions is configured to cause the processor to compare the plurality of target text strings to the plurality of ambiguous text strings in each database record by comparing each target text string to each ambiguous text string that is associated with the same locus as that target text string. 
     
     
       20. The non-transitory computer readable medium of  claim 18 , wherein the plurality of instructions is further configured to cause the processor to identify a matching database record in the database for which each of the plurality of target text strings matches one of the plurality of ambiguous text strings in the matching database record.

Join the waitlist — get patent alerts

Track US12183435B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.