US2014188396A1PendingUtilityA1

Oligomer sequences mapping

Assignee: COMPLETE GENOMICS INCPriority: Feb 3, 2009Filed: Dec 3, 2013Published: Jul 3, 2014
Est. expiryFeb 3, 2029(~2.5 yrs left)· nominal 20-yr term from priority
G16B 30/10G16B 30/20G16B 30/00G06F 19/22
65
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Mapping oligomer sequences includes receiving a set of related oligomer sequences, applying one or more key patterns derived from a set of oligomer sequence relationships to obtain one or more keys that are consistent with the set of related oligomer sequences, and locating the one or more keys in an index configured to map a plurality of possible keys to their respective candidate and/or validated locations in a reference.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of sequence mapping, the method comprising:
 receiving a data set of sequences corresponding to a plurality of fragments of genetic material, the data set including a first sequence, wherein the first sequence comprises at least one ambiguous position;   applying a key pattern to the first sequence to generate an initial key;   substituting one or more bases at the ambiguous position in the initial key to generate a set of one or more substituted keys from the initial key, wherein the substituted key is a possible match to a reference key in a reference index, the reference index being generated by respectively applying the key pattern to each of a plurality of locations of a reference sequence to obtain a plurality of reference keys, the application of the key pattern to each of the locations causing a selection of a same number of specified positions of the reference sequence relative to the respective location; and   comparing the substituted keys to the reference keys in the reference index to identify one or more reference keys that match to one or more of the substituted keys, thereby determining one or more candidate locations of the first sequence in the reference sequence, wherein the method is performed with a computer.   
     
     
         2 . The method of  claim 1 , wherein the first sequence has an unidentified base at the ambiguous position, and wherein the substituting including using the one or more bases at the ambiguous position. 
     
     
         3 . The method of  claim 1 , further comprising:
 validating a candidate location by comparing all bases of the first sequence to the indicated portions of the reference sequence.   
     
     
         4 . The method of  claim 3 , further comprising:
 outputting one or more validated locations in the reference sequence that match the first sequence.   
     
     
         5 . The method of  claim 1 , wherein the ambiguous position is substituted to create four possible keys to match to the reference index. 
     
     
         6 . The method of  claim 1 , wherein X n  bases are substituted at the at least one ambiguous position in the initial key to create the set of substituted keys, wherein n is a number of ambiguous positions in the initial key and X is the number of bases to be substituted at the ambiguous positions, n being an integer greater than or equal to one. 
     
     
         7 . The method of  claim 6 , wherein X is four. 
     
     
         8 . The method of  claim 6 , wherein n is greater than one, and wherein the n ambiguous positions are contiguous. 
     
     
         9 . The method of  claim 6 , wherein n is greater than one, and wherein the n ambiguous positions are not contiguous. 
     
     
         10 . The method of  claim 1 , wherein applying the key pattern to the first sequence includes reordering bases from the first sequence. 
     
     
         11 . The method of  claim 1 , wherein the initial key specifies a base at the ambiguous position, and wherein the set of substituted keys includes the initial key. 
     
     
         12 . The method of  claim 1 , wherein the set of substituted keys reflect all possible bases or combinations of bases at the at least one ambiguous position. 
     
     
         13 . The method of  claim 1 , wherein a majority or all of the positions in the first key or the data set are individually substituted to create the set of modified first keys. 
     
     
         14 . A computer program product for oligomer sequence mapping, the computer program product being embodied in a non-transitory computer readable medium and comprising computer instructions for:
 receiving a data set of sequences corresponding to a plurality of fragments of genetic material, the data set including a first sequence, wherein the first sequence comprises at least one ambiguous position;   applying a key pattern to the first sequence to generate an initial key;   substituting one or more bases at the ambiguous position in the initial key to generate a set of one or more substituted keys from the initial key, wherein the substituted key is a possible match to a reference key in a reference index, the reference index being generated by respectively applying the key pattern to each of a plurality of locations of a reference sequence to obtain a plurality of reference keys, the application of the key pattern to each of the locations causing a selection of a same number of specified positions of the reference sequence relative to the respective location; and   comparing the substituted keys to the reference keys in the reference index to identify one or more reference keys that match to one or more of the substituted keys, thereby determining one or more candidate locations of the first sequence in the reference sequence.   
     
     
         15 . The computer program product of  claim 14 , further comprising computer instructions for:
 validating a candidate location by comparing all bases of the first sequence to the indicated portions of the reference sequence.   
     
     
         16 . The computer program product of  claim 14 , wherein X n  bases are substituted at the at least one ambiguous position in the initial key to create the set of substituted keys, wherein n is a number of ambiguous positions in the initial key and X is the number of bases to be substituted at the ambiguous positions, n being an integer greater than or equal to one. 
     
     
         17 . The computer program product of  claim 14 , wherein the set of substituted keys reflect all possible bases or combinations of bases at the at least one ambiguous position. 
     
     
         18 . A system comprising for oligomer sequence mapping, the system comprising:
 a reference index generator configured to:
 receive a reference sequence and a key pattern; 
 generate a reference index by respectively applying the key pattern to each of a plurality of locations of a reference sequence to obtain a plurality of reference keys, the application of the key pattern to each of the locations causing a selection of a same number of specified positions of the reference sequence relative to the respective location; 
   a communication interference configured to receive a data set of sequences corresponding to a plurality of fragments of genetic material, the data set including a first sequence, wherein the first sequence comprises at least one ambiguous position;   a key generator configured to:
 apply a key pattern to the first sequence to generate an initial key; 
 substitute one or more bases at the ambiguous position in the initial key to generate a set of one or more substituted keys from the initial key, wherein the substituted key is a possible match to a reference key in the reference index; and 
   a reference index search engine configured to:
 compare the substituted keys to the reference keys in the reference index to identify one or more reference keys that match to one or more of the substituted keys, thereby determining one or more candidate locations of the first sequence in the reference sequence. 
   
     
     
         19 . The system of  claim 18 , wherein the reference index search engine is further configured to:
 validate a candidate location by comparing all bases of the first sequence to the indicated portions of the reference sequence.   
     
     
         20 . The system of  claim 18 , wherein X n  bases are substituted at the at least one ambiguous position in the initial key to create the set of substituted keys, wherein n is a number of ambiguous positions in the initial key and X is the number of bases to be substituted at the ambiguous positions, n being an integer greater than or equal to one. 
     
     
         21 . The system of  claim 18 , wherein the set of substituted keys reflect all possible bases or combinations of bases at the at least one ambiguous position.

Join the waitlist — get patent alerts

Track US2014188396A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.