US2008082273A1PendingUtilityA1

Computer algorithm for automatic allele determination from fluorometer genotyping device

Assignee: APPLERA CORPPriority: Mar 1, 2002Filed: Sep 5, 2007Published: Apr 3, 2008
Est. expiryMar 1, 2022(expired)· nominal 20-yr term from priority
G16B 40/00G06F 16/285G16B 20/00G16B 20/40G16B 20/20G16B 40/10
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present invention provides methods and systems for an automated method of identifying allele values from data files derived from processed fluorophore emissions detected during the observation of fluorophore labeled nucleotide probes used in analyzing polymorphic DNA are provided. These methods are used in the rapid and efficient distinguishing of targeted polymorphic DNA sites without control samples.

Claims

exact text as granted — not AI-modified
1 . A method for categorizing a dataset comprising a plurality of datapoints, each datapoint comprising at least two numerical values, said method comprising the steps of: 
 (a) producing a plurality of angular values by calculating an angular value for each datapoint based on said datapoint's numerical values;    (b) sorting said plurality of datapoints by said angular values;    (c) producing a plurality of difference values by calculating differences between adjacent angular values;    (d) determining at least one category-dividing value by identifying at least one difference value above a predetermined threshold gap value; and    (e) classifying at least one datapoint according to its angular value relative to at least one category-dividing value.    
     
     
         2 . The method of  claim 1  wherein each datapoint comprises two numerical values.  
     
     
         3 . The method of  claim 2  wherein said angular value is an arctangent of said two numerical values.  
     
     
         4 . The method of  claim 1  wherein said numerical values represent fluorometric data.  
     
     
         5 . The method of  claim 1  wherein said determining step (d) identifies two category-dividing values.  
     
     
         6 . The method of  claim 1  further comprising the step of normalizing said numerical values to a scale.  
     
     
         7 . The method of  claim 6  wherein said scale ranges from 0.0 to 1.0.  
     
     
         8 . The method of  claim 1  further comprising the step of removing non-amplification datapoints from said dataset, said step comprising the steps of: 
 (i) calculating a Euclidean distance for each datapoint;    (ii) removing at least one datapoint from said dataset, wherein the Euclidean distance of said datapoint falls below a predetermined distance threshold.    
     
     
         9 . The method of  claim 1  wherein said determining step (d) identifies two category-dividing values comprising a first and a second category-dividing value, and said classifying step (e) comprises the steps of: 
 (i) classifying at least one datapoint in a first category, wherein all datapoints of said first category have an angular value lower than said first and second category-dividing values;    (ii) classifying at least one datapoint in a second category, wherein all datapoints of said second category have an angular value between said first and second category-dividing values; and    (iii) classifying at least one datapoint in a third category, wherein all datapoints of said third category have an angular value greater than said first and second category-dividing values.    
     
     
         10 . The method of  claim 9  wherein classification in said first category corresponds to homozygosity for a first allele, classification in said third category corresponds to homozygosity for a second allele, and classification in said second category corresponds to heterozygosity for said first and second alleles.  
     
     
         11 . The method of  claim 10  further comprising the step of determining the presence of a condition to bring to the attention of a human user, wherein said condition comprises the proportion of datapoints classified as heterozygous exceeding a predetermined threshold.  
     
     
         12 . The method of  claim 11  further comprising the step of determining the presence of a condition to bring to the attention of a human user.  
     
     
         13 . The method of  claim 12  wherein said condition comprises a substantial majority of datapoints being classified in one category.  
     
     
         14 . The method of  claim 13  wherein said category corresponds to heterozygosity for a first and second allele.  
     
     
         15 . The method of  claim 13  wherein said category corresponds to homozygosity for either a first or second allele.  
     
     
         16 . The method of  claim 13  wherein said category cannot be determined to correspond to either heterozygosity or homozygosity.  
     
     
         17 . The method of  claim 12  wherein said condition comprises said datapoints being classified into more than three categories.  
     
     
         18 . The method of  claim 12  wherein said condition comprises at least one of said datapoints remaining unclassified.  
     
     
         19 . The method of  claim 12  wherein said condition comprises the Euclidean distance between at least one of said classified datapoints and at least one non-amplification datapoint being below a predetermined threshold.  
     
     
         20 . The method of  claim 12  wherein said condition comprises a substantial majority of datapoints in said first category having an angular value higher than a predetermined threshold.  
     
     
         21 . The method of  claim 20  wherein said angular value is an arctangent and said predetermined threshold is 0.67.  
     
     
         22 . The method of  claim 12  wherein said condition comprises a substantial majority of datapoints in said third category having an angular value lower than a predetermined threshold.  
     
     
         23 . The method of  claim 22  wherein said angular value is an arctangent and said predetermined threshold is 1.0.  
     
     
         24 . The method of  claim 12  wherein said condition comprises a substantial majority of datapoints in said second category having an angular value lower than a first predetermined threshold or higher than a second predetermined threshold.  
     
     
         25 . The method of  claim 24  wherein said angular value is an arctangent, said first predetermined threshold is 0.18, and said second predetermined threshold is 1.35.  
     
     
         26 . The method of  claim 12  wherein said condition comprises the difference between the largest angular value of a datapoint in a category and the smallest angular value of a datapoint in the category exceeding a predetermined threshold.  
     
     
         27 . The method of  claim 26  wherein said angular value is an arctangent and said second predetermined threshold is 0.6.  
     
     
         28 . The method of  claim 12  wherein said first allele is a major allele and said second allele is a minor allele, and said major and minor alleles are in a Hardy-Weinberg equilibrium.  
     
     
         29 . The method of  claim 28  further comprising the step of determining the presence of a condition to bring to the attention of a human user, wherein said condition indicates an incompatibility with a Hardy-Weinberg equilibrium.  
     
     
         30 . The method of  claim 29  wherein said incompatibility comprises a greater number of datapoints classified as homozygous for said minor allele than classified as heterozygous.  
     
     
         31 . The method of  claim 12  further comprising the step of determining the presence of a condition to bring to the attention of a human user, said determining step comprising the steps of: (i) calculating the center of the set of removed datapoints, said center comprising an x and y coordinate; and (ii) determining if either said x or y coordinate exceeds a predetermined threshold.  
     
     
         32 . The method of  claim 31  wherein said predetermined threshold is 0.3 on a normalized scale of 0.0 to 1.0.  
     
     
         33 . A method for performing allelic differentiation comprising: 
 acquiring fluorescence intensity data for a plurality of samples wherein the fluorescence intensity data is obtained by amplification of each sample in the presence of at least two fluorophore labels;    generating an angular value for each sample by comparing the fluorescence intensity obtained for the at least tow fluorophore labels;    arranging the samples according to their angular value to form an angular-valued based distribution of the samples;    determining a difference value for each sample by taking the difference between the angular value for a selected sample and the angular value for an adjacent sample;    associating at least one difference value range with a selected allelic composition;    evaluating each sample's difference value with respect to the at least one difference value range to determine if the sample resides within the range; and    identifying the allelic composition of each sample on the basis of the difference value range which the sample resides within.    
     
     
         34 . The method of  claim 33 , wherein the allelic compositions comprises a homozygous allele.  
     
     
         35 . The method of  claim 33 , wherein the allelic compositions comprises a heterozygous allele.  
     
     
         36 . The method of  claim 33 , wherein the angular values for each sample are normalized.  
     
     
         37 . The method of  claim 33 , wherein the angular values are calculated as the arctangent between the at least two fluorophore labels.  
     
     
         38 . The method of  claim 33 , further comprising reducing the number of samples undergoing analysis by: 
 calculating a Euclidean distance for each sample; and    identifying a Euclidean distance threshold for which samples having a Euclidean distance below the threshold are removed from further analysis.    
     
     
         39 . A method for genotypic analysis comprising: 
 amplifying a plurality of a genetic samples in the presence of at least two discriminable labels to thereby obtain intensity information indicative of the signals generated by the at least tow discriminable labels during amplification;    calculating an angular value for each sample by comparing the intensity information for the at least two discriminable labels used during amplification;    ordering the samples on the basis of their angular value;    calculating a difference value for each sample by taking the difference between the angular value for a selected sample and the angular value for an adjacent sample;    identifying difference value ranges corresponding to homozygous and heterozygous allelic variations; and    determining whether a sample corresponds to a homozygous or heterozygous allelic variant by determining if the sample's difference value resides within the difference value ranges corresponding to homozygous or heterozygous allelic variation.    
     
     
         40 . The method of  claim 39 , wherein the angular values for each sample normalized prior to difference value determination.  
     
     
         41 . The method of  claim 39 , wherein the angular values are calculated as the arctangent between the at least two discriminable labels.  
     
     
         42 . The method of  claim 41 , further comprising identifying at least one of the discriminable labels as undetermined based on at least one predetermined condition.  
     
     
         43 . The method of  claim 42 , wherein the at least one predetermined condition is based on comparison of the dye emissions to a control probe.  
     
     
         44 . The method of  claim 43 , wherein the at least one predetermined condition is based on a range between a maximum and minimum fluorescence dye emission.  
     
     
         45 . A method for genotypic analysis comprising: 
 receiving intensity information for a plurality of samples comprising intensity data of at least first and second fluorescent probes;    calculating a range between a maximum and minimum of the emission data for each of the first and second fluorescent probes;    determining whether the range between the maximum and minimum of the emission data for each of the at least first and second fluorescent probes exceeds a predetermined threshold;    determining whether a plurality of clusters of datapoints defined by emission data of the at least first and second fluorescent probes is present based on whether the range between the maximum and minimum of the emission data exceeds the predetermined threshold for each of the at least first and second fluorescent probes; and,    when a plurality of clusters are present    calculating an angular value for each sample of the plurality of sample by comparing the intensity information for the at least first and second fluorescent,    ordering the plurality of samples on the basis of their angular value,    calculating a difference value for each sample by taking the difference between the angular value for a selected sample and the angular value for an adjacent sample,    identifying difference value ranges corresponding to homozygous and heterozygous allelic variations, and    determining whether a sample corresponds to a homozygous or heterozygous allelic variant by determining if the sample's difference value resides within the difference value ranges corresponding to homozygous or heterozygous allelic variation.

Join the waitlist — get patent alerts

Track US2008082273A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.