US2013261005A1PendingUtilityA1

System and methods for indel identification using short read sequencing

Assignee: APPLIED BIOSYSTEMS LLCPriority: Feb 5, 2007Filed: May 30, 2013Published: Oct 3, 2013
Est. expiryFeb 5, 2027(~0.5 yrs left)· nominal 20-yr term from priority
Inventors:Zheng Zhang
G16B 30/10G16B 30/20G16B 30/00G06F 19/22
74
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems, methods, and analytical approaches for short read sequence assembly and for the detection of insertions and deletions (indels) in a reference genome. A method suitable for software implementation is presented in which indels may be readily identified in a computationally efficient manner.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . The method of  claim 22 , wherein the second mapping operation further identifies indels with respect to the reference sequence by determining a difference between an expected intervening sequence length between the non-overlapping sequences of a mate pair and an observed intervening sequence length between the non-overlapping sequences of a mate pair. 
     
     
         3 . The method of  claim 2 , wherein the indel comprises an insertion occurring between the non-overlapping sequences of a mate pair which accounts for the difference between the expected intervening sequence length and the observed intervening sequence length. 
     
     
         4 . The method of  claim 2 , wherein the indel comprises a deletion occurring between the non-overlapping sequences of a mate pair which accounts for the difference between the expected intervening sequence length and the observed intervening sequence length. 
     
     
         5 . The method of  claim 22 , wherein the first nucleic acid sequence information comprises paired read sequence information separated by the intervening sequence whose length is within a known range. 
     
     
         6 . The method of  claim 5 , wherein each of the paired read sequences has a length of between approximately 10 and 75 bases. 
     
     
         7 . The method of  claim 5 , wherein the intervening sequence has a length of between approximately 2 kilobases and 15 kilobases. 
     
     
         8 .- 20 . (canceled) 
     
     
         21 . A method of nucleic acid sequence analysis, comprising:
 sequencing a first locality from a plurality of nucleic acid fragments to generate first non-overlapping sequences of a plurality of mate pair sequences;   sequencing a second locality from the nucleic acid fragments to generate second non-overlapping sequences of the mate pair sequences, the first and second locality of a nucleic acid fragment corresponds to a first region of a genome and a second region of a genome, the first and second regions separated by an intervening region whose length is within a known range;   receiving second nucleic acid sequence information comprising at least one reference sequence;   performing a computer assisted mapping operation for the mate pair sequences in which the first non-overlapping pairwise sequence and the second non-overlapping pairwise sequence for a respective mate pair are aligned to the at least one reference sequence using a processor by the steps of:
 performing a first mapping operation using a processor to align the first non-overlapping pairwise sequence of the mate pair sequences to the at least one reference sequence with a first selected mismatch constraint, 
 identifying mate pair sequences having first non-overlapping pairwise sequences which are aligned to the at least one reference sequence while satisfying the selected mismatch constraint, 
 designating a window region within the at least one reference sequence for the identified mate pair sequences based on the alignment of the first non-overlapping pairwise sequence to the at least one reference sequence, 
 performing a second mapping operation using a processor to align the second non-overlapping pairwise sequence to the window region of the reference sequence with a second selected mismatch constraint, 
 identifying mate pair sequences with first and second non-overlapping pairwise sequences that have mapped to the at least one reference sequence following performing the first and second mapping operations; 
   and,   outputting the results of the mapping operations.   
     
     
         22 . A method for nucleic acid sequence analysis, comprising:
 sequencing a first locality and a second locality from at least one nucleic acid fragment to generate first and second non-overlapping sequences, the first locality corresponding to a first region of a genome and the second locality corresponding to a second region of a genome, the first and second regions separated by an intervening region;   receiving a reference nucleic acid sequence information comprising at least one reference sequence;   performing a first mapping operation using a processor to at least partially align the first non-overlapping sequence to the reference nucleic acid sequence,   designating a window region comprising at least a portion of the at least one reference sequence based at least in part upon the alignment of the first non-overlapping sequence to the at least one reference sequence,   performing a second mapping operation using a processor to at least partially align the second non-overlapping sequence to the designated window region of the reference sequence thereby positioning the first and second non-overlapping sequences with respect to one another, and   outputting the results of the mapping operations.   
     
     
         23 . A method for nucleic acid sequence analysis, comprising:
 sequencing at least a portion of a plurality nucleic acid fragments to generate a plurality of sequence reads;   receiving a reference nucleic acid sequence information comprising at least one reference sequence;   mapping the plurality of sequence reads using a processor to at least partially align the sequence reads to the reference nucleic acid sequence;   identifying a first sequence read that partially maps to a first portion of the reference nucleic acid sequence and second sequence read that partially maps to a second portion of the reference nucleic acid sequence;   traversing a shortest path along overlapping sequence reads using a processor to determine a sequence of an indel between the first portion of the reference nucleic acid sequence and the second portion of the reference nucleic acid sequence; and   outputting the sequence of the indel.   
     
     
         24 . The method of  claim 23 , further comprising identifying overlapping sequence reads comprising a sequence read that partially overlaps with the first sequence read and sequence read that overlaps the second sequence read. 
     
     
         25 . The method of  claim 23 , wherein overlapping reads are determined by a match of a minimum number of bases of an overlap region between reads.

Join the waitlist — get patent alerts

Track US2013261005A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.