US2010094563A1PendingUtilityA1

System and Method for Consensus-Calling with Per-Base Quality Values for Sample Assemblies

Assignee: APPLIED BIOSYSTEMSPriority: Oct 25, 2001Filed: Jul 29, 2008Published: Apr 15, 2010
Est. expiryOct 25, 2021(expired)· nominal 20-yr term from priority
Inventors:Jon Sorenson
G16B 30/10G16B 30/00
68
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present teachings disclose a method for evaluation of a polynucleotide sequence using a consensus-based analysis approach. The sequence analysis method utilizes quality values for a plurality of aligned sequence fragments to identify consensus basecalls and calculate associated consensus quality values. The disclosed method is applicable to resolution of single nucleotide polymorphisms, mixed-based sequences, heterozygous allelic variants, and heterogeneous polynucleotide samples.

Claims

exact text as granted — not AI-modified
1 . A basecalling method for predicting the composition of a polynucleotide sequence, the method comprising:
 receiving information for a plurality of sequence fragments comprising basecalls spanning at least a portion of the sample sequence and corresponding quality values indicative of a calculated degree of confidence in the basecalls;   aligning the plurality of sequence fragments to identify regions of basecall overlap between the sequence fragments; and   calculating a consensus basecall and a consensus quality value for the regions of basecalling overlap in the aligned sequence fragments wherein the consensus quality values are determined using the quality values for the basecalls of the sequence fragments.   
     
     
         2 . The basecalling method of  claim 1 , wherein mixed-basecalls present in the sequence fragments are evaluated during calculation of the consensus basecall by comparing the basecalls in the regions of basecall overlap to identify one or more constituent pure basecalls. 
     
     
         3 . The basecalling method of  claim 1 , wherein calculation of the consensus basecall further comprises distinguishing a noise-related component for at least one of the mixed-basecalls to thereby resolve the mixed-basecall into one or more constituent pure basecalls. 
     
     
         4 . The basecalling method of  claim 1 , wherein the consensus basecall is calculated by applying a differential weight to each of the quality values for the basecalls of the aligned sequence fragments and thereafter the consensus basecall is assigned as the basecall for the sequence fragments with the greatest overall quality value. 
     
     
         5 . The basecalling method of  claim 4 , wherein calculation of the consensus basecall further comprises identifying the intersection between one or more of the basecalls for the sequence fragments. 
     
     
         6 . The basecalling method of  claim 5 , wherein identification of the intersection between one or more of the basecalls for the sequence fragments is used to resolve mixed-basecalls into one or more constituent pure basecalls. 
     
     
         7 . The basecalling method of  claim 6 , wherein when no intersection between one or more of the basecalls for the sequence fragments is observed then the consensus basecall is assigned as one of the basecalls for the sequence fragments with a reduced quality value. 
     
     
         8 . The basecalling method of  claim 1 , wherein the consensus basecall is determined by identifying an agreeing basecall for the aligned sequence fragments and the consensus basecall is assigned as the agreeing basecall. 
     
     
         9 . The basecalling method of  claim 8 , wherein agreeing basecall identification is performed following re-calling of one or more of the basecalls for the sequence fragments using a more stringent basecalling criteria. 
     
     
         10 . The basecalling method of  claim 8 , wherein when a lack of agreement in the basecalls for the aligned sequence fragments is observed then the consensus basecall is determined using a weighted basecall vote. 
     
     
         11 . The basecalling method of  claim 1 , wherein calculation of the consensus basecall is used to identify heterozygosity between allelic variants of the polynucleotide sequence. 
     
     
         12 . The basecalling method of  claim 1 , wherein calculation of the consensus basecall is used to identify single nucleotide polymorphisms contained within the polynucleotide sequence. 
     
     
         13 . The basecalling method of  claim 1 , wherein calculation of the consensus basecall is used to identify heterogeneous polynucleotide sequence populations. 
     
     
         14 . A system for predicting the composition of a polynucleotide sequence, comprising:
 a sample processing module that receives information for a plurality of sequence fragments, the sample processing module providing functionality for identifying basecalls spanning at least a portion of the sample sequence and corresponding quality values indicative of a calculated degree of confidence in the basecalls;   a specimen processing module that assembles the plurality of sequence fragments to identify regions of basecall overlap between the sequence fragments; and   a project processing module that calculates a consensus basecall and a consensus quality value for the regions of basecalling overlap in the assembled sequence fragments wherein the consensus quality values are determined using the quality values for the basecalls of the sequence fragments.   
     
     
         15 . The system of  claim 14 , wherein the project processing module evaluates mixed-basecalls present in the sequence fragments during calculation of the consensus basecall. 
     
     
         16 . The system  claim 14 , wherein the project processing module further identifies the consensus basecall by distinguishing a noise-related component for at least one of the mixed-basecalls to thereby resolve the mixed-basecall into one or more constituent pure basecalls. 
     
     
         17 . The system  claim 14 , wherein the project processing module calculates the consensus basecall by applying a differential weight to each of the quality values for the basecalls of the aligned sequence fragments and thereafter assigns the consensus basecall as the basecall for the sequence fragments with the greatest overall quality value. 
     
     
         18 . The system of  claim 17 , wherein the project processing module calculates the consensus basecall by identifying the intersection between one or more of the basecalls for the sequence fragments. 
     
     
         19 . The system of  claim 18 , wherein the project processing module identifies the intersection between one or more of the basecalls for the sequence fragments to resolve mixed-basecalls into one or more constituent pure basecalls. 
     
     
         20 . The system of  claim 14 , wherein the project processing module determines the consensus basecall by identifying an agreeing basecall for the aligned sequence fragments and thereafter assigns the consensus basecall as the agreeing basecall. 
     
     
         21 . The system of  claim 14 , wherein the modules are used to identify heterozygosity between allelic variants of the polynucleotide sequence. 
     
     
         22 . The system of  claim 14 , wherein the modules are used to identify single nucleotide polymorphisms contained within the polynucleotide sequence. 
     
     
         23 . The system of  claim 14 , wherein the modules are used to identify heterogeneous polynucleotide sequence populations. 
     
     
         24 . The system of  claim 14 , wherein the modules comprise a single unified sequence analysis module. 
     
     
         25 . The system of  claim 14 , wherein the modules are integrated into a sequence analysis software application.

Join the waitlist — get patent alerts

Track US2010094563A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.