US2023410946A1PendingUtilityA1

Systems and methods for sequence data alignment quality assessment

Assignee: LIFE TECHNOLOGIES CORPPriority: Jul 6, 2010Filed: Jun 21, 2023Published: Dec 21, 2023
Est. expiryJul 6, 2030(~3.9 yrs left)· nominal 20-yr term from priority
G16B 30/10G16B 40/00G16B 30/00G16B 30/20G06N 7/01G06N 3/126
83
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A computer-implemented method for classifying alignments of paired nucleic acid sequence reads is disclosed. A plurality of paired nucleic acid sequence reads is received, wherein each read is comprised of a first tag and a second tag separated by an insert region. Potential alignments for the first and second tags of each read to a reference sequence is determined, wherein the potential alignments satisfies a minimum threshold mismatch constraint. Potential paired alignments of the first and second tags of each read are identified, wherein a distance between the first and second tags of each potential paired alignment is within an estimated insert size range. An alignment score is calculated for each potential paired alignment based on a distance between the first and second tags and a total number of mismatches for each tag.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled) 
     
     
         21 . A method for classifying alignments of paired nucleic acid sequence reads, comprising:
 receiving, by a computing device comprising a processor and memory, a plurality of paired nucleic acid sequence reads, wherein each paired nucleic acid sequence read comprises a first read of a first tag derived from a first region of a polynucleotide and a second read of a second tag derived from a second region of the polynucleotide, the first and second reads produced by next generation sequencing, wherein the first tag and the second tag are separated by an insert region;   mapping, by the computing device, the first and second reads of each paired nucleic acid sequence read to a reference genome to form potential alignments, wherein one or more potential alignments are produced for each paired nucleic acid sequence read, wherein each potential alignment satisfies a minimum threshold mismatch constraint;   identifying, by the computing device, potential paired alignments of the first and second reads of each paired nucleic acid sequence read, wherein a distance between the first and second reads of each potential paired alignment is within an estimated insert size range;   calculating, by the computing device, an alignment score for each potential paired alignment based on:
 a distance between the first and second reads, and 
 a total number of mismatches for the first and second reads; and 
   selecting, by the computing device, the potential paired alignment having a highest alignment score as a primary alignment for each paired nucleic acid sequence read.   
     
     
         22 . The method, as recited in  claim 21 , wherein the paired nucleic acid sequence read is a mate-pair read. 
     
     
         23 . The method, as recited in  claim 21 , wherein the paired nucleic acid sequence read is a paired-end read. 
     
     
         24 . The method, as recited in  claim 21 , wherein the estimated insert size range is based on a standard deviation value derived from a distribution of estimated sizes of insert region for the plurality of paired nucleic acid sequence reads. 
     
     
         25 . The method, as recited in  claim 21 , wherein the calculated alignment score is a function of read alignment length. 
     
     
         26 . The method, as recited in  claim 21 , wherein the calculated alignment score is a function of a total number of possible alignments for each read. 
     
     
         27 . A system for sequence alignment quality assessment, comprising:
 a next generation sequencing instrument configured to interrogate a sample to produce a plurality of read sequences from the sample; and   a processor in communication with the next generation sequencing instrument, the processor configured to,
 obtain the read sequences from the next generation sequencing instrument, 
 map the read sequences to a reference genome to produce potential alignments, wherein one or more potential alignments are produced for a given read sequence, wherein each potential alignment satisfies a minimum threshold mismatch constraint, 
 calculate a quality value for each potential alignment, 
 output the potential alignments and associated quality values, and 
 select the potential alignment having a highest quality value as a primary alignment for the given read sequence. 
   
     
     
         28 . The system as recited in  claim 27 , wherein the quality value is calculated for the potential alignment of the read sequence corresponding to a single fragment library type. 
     
     
         29 . The system as recited in  claim 27 , wherein the quality value is calculated for the potential alignment of the read sequence corresponding to a paired read library type. 
     
     
         30 . The system as recited in  claim 29 , wherein for the paired read library type, the mapping produces aligned paired reads, wherein the aligned paired reads must have insert region sizes that fall within an estimated insert size range for the aligned paired reads, wherein the aligned paired reads are separated by an insert region. 
     
     
         31 . The system as recited in  claim 30 , wherein the estimated insert size range is based on a standard deviation value derived from a distribution of estimated insert region sizes of the aligned paired reads. 
     
     
         32 . A method for sequence alignment quality assessment, comprising:
 interrogating a sample, by a next generation sequencing instrument, to produce a plurality of read sequences from the sample;   obtaining, by a computing device comprising a processor and a memory, the plurality of read sequences from the next generation sequencing instrument;   mapping, by the computing device, the read sequences to a reference genome to produce potential alignments, wherein one or more potential alignments are produced for a given read sequence, wherein each potential alignment satisfies a minimum threshold mismatch constraint;   calculating, by the computing device, a quality value for each potential alignment;   outputting, by the computing device, the potential alignments and associated quality values; and   selecting, by the computing device, the potential alignment having a highest quality value as a primary alignment.   
     
     
         33 . The method, as recited in  claim 32 , wherein the quality value is calculated for the potential alignment of the read sequence corresponding to a single fragment library type. 
     
     
         34 . The method, as recited in  claim 32 , wherein the quality value is calculated for the potential alignment of the read sequence corresponding to a paired read library type. 
     
     
         35 . The method, as recited in  claim 32 , wherein the calculated quality value for each potential alignment is a function of read alignment length. 
     
     
         36 . The method, as recited in  claim 32 , wherein the calculated quality value for each potential alignment is a function of number of read mismatches. 
     
     
         37 . The method, as recited in  claim 34 , wherein for the paired read library type, the mapping produces aligned paired reads, wherein the aligned paired reads must have insert region sizes that fall within an estimated insert size range for the aligned paired reads, wherein the aligned paired reads are separated by an insert region. 
     
     
         38 . The method, as recited in  claim 37 , wherein the estimated insert size range is based on a standard deviation value derived from a distribution of estimated insert region sizes of the aligned paired reads.

Join the waitlist — get patent alerts

Track US2023410946A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.