US2011270533A1PendingUtilityA1

Systems and methods for analyzing nucleic acid sequences

Assignee: LIFE TECHNOLOGIES CORPPriority: Apr 30, 2010Filed: Jul 13, 2011Published: Nov 3, 2011
Est. expiryApr 30, 2030(~3.8 yrs left)· nominal 20-yr term from priority
G16B 30/00G16B 30/10G16B 30/20
55
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Nucleic acid sequence mapping/assembly methods are disclosed. The methods initially map only a contiguous portion of each read to a reference sequence and then extends the mapping of the read at both ends of the mapped contiguous portion until the entire read is mapped (aligned). In various embodiments, a mapping score can be calculated for the read alignment using a scoring function, score (i, j)=M+mx, where M can be the number of matches in the extended alignment, x can be the number of mismatches in the alignment, and m can be a negative penalty for each mismatch. The mapping score can be utilized to rank or choose the best alignment for each read.

Claims

exact text as granted — not AI-modified
1 . A system for mapping a nucleic acid sequence read to a reference sequence, comprising:
 a first data store configured to store nucleic acid sequencing data;   a second data store configured to store reference sequence data; and   a computing device in communication with the first data store and the second data store, comprising:
 a mapping module configured to:
 obtain the nucleic acid sequence read from the first data store, 
 obtain a reference sequence from the second data store, 
 select a contiguous portion of the nucleic acid sequence read, 
 map the contiguous portion of the nucleic acid sequence read to the reference sequence using an approximate string mapping method that allows for a set number of mismatches between the contiguous portion and the reference sequence and produces at least one match of the contiguous portion to the reference sequence, and 
 map a remaining portion of the nucleic acid sequence read to the reference sequence using an ungapped local alignment method that produces an alignment of the remaining portion extending from the at least one match to complete a map of the nucleic acid sequence read to the reference sequence. 
 
   
     
     
         2 . The system of  claim 1 , wherein the mapping module is further configured to automatically select a contiguous portion of the nucleic acid sequence read and map the contiguous portion to the reference sequence iteratively. 
     
     
         3 . The system of  claim 2 , wherein the mapping module is further configured to automatically select a contiguous portion at a different location on the nucleic acid sequence read but with the same length on the read at each iteration until at least one match is produced. 
     
     
         4 . The system of  claim 2 , wherein the mapping module is further configured to select a contiguous portion at a same location on the nucleic acid sequence read but with a different length on the read at each iteration until a number of matches of the contiguous portion to the reference sequence is less than a certain threshold number. 
     
     
         5 . The system of  claim 1 , wherein the alignment extends from the at least one match in either direction. 
     
     
         6 . The system of  claim 1 , wherein the ungapped local alignment method uses a scoring function to select an alignment with the best score. 
     
     
         7 . The system of  claim 6 , wherein the scoring function is a sum of a number of matches in the alignment and a product of a number of mismatches in the alignment and a negative penalty for each mismatch. 
     
     
         8 . The system of  claim 1 , wherein the number of mismatches is user selected. 
     
     
         9 . The system of  claim 4 , wherein the threshold number is user selected. 
     
     
         10 . The system of  claim 1 , wherein the contiguous portion is user selected. 
     
     
         11 . The system of  claim 1 , wherein the first data store and the second data store is hosted on a single data storage device. 
     
     
         12 . The system of  claim 1 , wherein either the first data store or the second data store is hosted on the computing device. 
     
     
         13 . The system of  claim 1 , wherein both the first data store and the second data store are hosted on the computing device 
     
     
         14 . The system of  claim 1 , wherein the first data store is hosted by a nucleic acid sequencer. 
     
     
         15 . A computer implemented method for mapping a nucleic acid sequence read to a reference sequence, comprising:
 obtaining the nucleic acid sequence read;   obtaining a reference sequence;   selecting a contiguous portion of the nucleic acid sequence read;   mapping the contiguous portion of the nucleic acid sequence read to the reference sequence using an approximate string mapping method that allows for a set number of mismatches between the contiguous portion and the reference sequence and produces at least one match of the contiguous portion to the reference sequence; and   mapping a remaining portion of the nucleic acid sequence read to the reference sequence using an ungapped local alignment method that produces an alignment of the remaining portion extending from the at least one match to complete a map of the nucleic acid sequence read to the reference sequence.   
     
     
         16 . The method of  claim 15 , further comprising iteratively selecting a contiguous portion of the nucleic acid sequence read and mapping the contiguous portion to the reference sequence. 
     
     
         17 . The method of  claim 16 , wherein iteratively selecting a contiguous portion of the nucleic acid sequence read comprises selecting a contiguous portion at a different location but with the same length on the read at each iteration until the at least one match is produced. 
     
     
         18 . The method of  claim 16 , wherein iteratively selecting a contiguous portion of the nucleic acid sequence read comprises selecting a contiguous portion at a same location but with a different length on the read at each iteration until a number of matches of the contiguous portion to the reference sequence is less than a certain threshold number. 
     
     
         19 . The method of  claim 15 , wherein the alignment extends from the at least one match in either direction. 
     
     
         20 . The method of  claim 15 , wherein the ungapped local alignment method uses a scoring function and selects an alignment with the best score. 
     
     
         21 . The method of  claim 20 , wherein the scoring function is a sum of a number of matches in the alignment and a product of a number of mismatches in the alignment and a negative penalty for each mismatch. 
     
     
         22 . A non-transitory computer-readable storage medium whose contents include a program with instructions being executed on a processor so as to perform a method for mapping a read of a nucleic acid sequence to a reference sequence, the method comprising:
 providing a system, wherein the system comprises one or more distinct software module configured to:   receiving a nucleic acid sequence read;   obtaining a reference sequence;   selecting a contiguous portion of the nucleic acid sequence read;   mapping the contiguous portion to the reference sequence using an approximate string mapping method that allows for a set number of mismatches between the contiguous portion and the reference sequence and produces at least one match of the contiguous portion to the reference sequence; and   mapping a remaining portion of the read to the reference sequence using an ungapped local alignment method that produces an alignment of the remaining portion extending from the at least one match to complete a map of the nucleic acid sequence read to the reference genome.   
     
     
         23 . The non-transitory computer-readable storage medium of  claim 22 , further comprising iteratively selecting a contiguous portion of the read and mapping the contiguous portion to the reference sequence. 
     
     
         24 . The non-transitory computer-readable storage medium of  claim 23 , wherein iteratively selecting a contiguous portion of the read comprises selecting a contiguous portion at a different location but with the same length on the read at each iteration until the at least one match is produced. 
     
     
         25 . The non-transitory computer-readable storage medium of  claim 23 , wherein iteratively selecting a contiguous portion of the read comprises selecting a contiguous portion at a same location but with a different length on the read at each iteration until a number of matches of the contiguous portion to the reference sequence is less than a certain threshold number. 
     
     
         26 . The non-transitory computer-readable storage medium of  claim 22 , wherein the alignment extends from the at least one match in either direction. 
     
     
         27 . The non-transitory computer-readable storage medium of  claim 22 , wherein the ungapped local alignment method uses a scoring function and selects an alignment with the best score. 
     
     
         28 . The non-transitory computer-readable storage medium of  claim 27 , wherein the scoring function is a sum of a number of matches in the alignment and a product of a number of mismatches in the alignment and a negative penalty for each mismatch. 
     
     
         29 . A system for assembling a nucleic acid sequence from a plurality of paired sequence reads, comprising:
 a first data store configured to store the plurality of paired sequence reads, wherein each of the plurality of paired sequence reads comprise a pair of pairwise sequences separated by an intervening length; and   a computing device in communications with the first data store, comprising:
 a first assembly module configured to assemble the plurality of paired sequence reads to form a scaffold of the nucleic acid sequence, wherein the scaffold comprise a plurality of contiguous sequences separated by a gap region, wherein each of the plurality of contiguous sequences is comprised of two or more paired sequence reads, and 
 a second assembly module configured to assemble hanging pairwise sequences of the assembled paired sequence reads to fill in the gap region.

Join the waitlist — get patent alerts

Track US2011270533A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.