US2022093207A1PendingUtilityA1

Genetic Copy Number Alteration Classifications

Assignee: SEQUENOM INCPriority: Jul 27, 2016Filed: Dec 7, 2021Published: Mar 24, 2022
Est. expiryJul 27, 2036(~10 yrs left)· nominal 20-yr term from priority
G06N 7/01G16B 20/20G16B 20/10G16B 20/00G16B 40/00G06N 7/005
66
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Technology provided herein relates in part to non-invasive classification of one or more genetic copy number alterations (CNAs) for a test sample. Certain methods include sampling a quantification of sequence reads from parts of a genome, generating a confidence determination, and using the confidence determination to enhance classification. Technology provided herein is useful for classifying a genetic CNA for a sample as part of non-invasive pre-natal (NIPT) testing and oncology testing, for example.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 selecting, by a computing device, a set of genomic portions each coupled to a sequence read quantification for a test sample, wherein:
 the set of genomic portions comprise portions of a reference genome to which sequence reads obtained for sample nucleic acid from the test subject have been mapped, 
 the set of genomic portions comprise (i) genomic portions within a candidate region identified as a candidate copy number variation region, and/or (ii) genomic portions outside the candidate region and within a selected region larger than the candidate region that includes the candidate region, and 
 the candidate region is a chromosome or sub-chromosome region; 
   shifting, by the computing device, a quantification of the sequence reads mapped to the genomic portions for the candidate region at or near a quantification of sequence reads mapped to genomic portions outside the candidate region and within the selected region;   sampling, by the computing device, the genomic portions from the selected region, wherein a number of sampled genomic portions is about equal to a number of genomic portions in the candidate region, thereby generating a sampled region, and repeating the sampling to generate a plurality of sampled regions;   normalizing, by the computing device, a sequence read quantification for each of the sampled regions, thereby generating a normalized quantification for each of the sampled regions;   generating, by the computing device, a confidence determination for the candidate region by a process comprising comparing each normalized quantification to a measure of variability of sequence read quantifications for the candidate region for reference samples not having a significant copy number variation in the candidate region; and   providing a classification for presence or absence of the copy number variation for the candidate region for the test sample according to the confidence determination.   
     
     
         2 . The method of  claim 1 , further comprising:
 sequencing the sample nucleic acid by a sequencer that generates the sequence reads, wherein the sample nucleic acid comprises circulating cell-free nucleic acid from the test sample of a pregnant female bearing a fetus;   mapping, by the computing device, the sequence reads to genomic portions of a reference genome; and   counting, by the computing device, the sequence reads mapped to the genomic portions, wherein the counting generates the quantification of the sequence reads mapped to the genomic portions of the reference genome.   
     
     
         3 . The method of  claim 1 , wherein the shifting further comprises:
 determining a median read count of genomic portions in the candidate region, thereby providing a candidate region median count;   determining a median read count of genomic portions outside the candidate region and in the selected region, thereby providing an outside median count;   subtracting the outside median count from the candidate region median count, thereby determining a shift.   
     
     
         4 . The method of  claim 1 , wherein each of the sequence read count quantifications for the candidate region for each of the reference samples is a read count fraction for each candidate region in each of the reference samples. 
     
     
         5 . The method of  claim 1 , wherein the quantification of sequence reads for genomic portions in the candidate region, outside region and/or selected region is a normalized quantification generated by a normalization process that normalizes GC bias or other bias. 
     
     
         6 . The method of  claim 1 , wherein the generating the confidence determination further comprises comparing each normalized quantification for each sampled region, or absolute value of each normalized quantification for each sampled region, to a measure of variability threshold for the reference samples. 
     
     
         7 . The method of  claim 1 , wherein the generating the confidence determination further comprises determining a proportion of the sampled regions for which the normalized quantification is greater than or less than the threshold. 
     
     
         8 . The method of  claim 7 , wherein the proportion is a ratio of (i) the number of sampled regions for which the normalized quantification is greater than or less than the threshold to (ii) the total number of sampled regions. 
     
     
         9 . A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform actions including:
 selecting a set of genomic portions each coupled to a sequence read quantification for a test sample, wherein:
 the set of genomic portions comprise portions of a reference genome to which sequence reads obtained for sample nucleic acid from the test subject have been mapped, 
 the set of genomic portions comprise (i) genomic portions within a candidate region identified as a candidate copy number variation region, and/or (ii) genomic portions outside the candidate region and within a selected region larger than the candidate region that includes the candidate region, and 
 the candidate region is a chromosome or sub-chromosome region; 
   shifting a quantification of the sequence reads mapped to the genomic portions for the candidate region at or near a quantification of sequence reads mapped to genomic portions outside the candidate region and within the selected region;   sampling the genomic portions from the selected region, wherein a number of sampled genomic portions is about equal to a number of genomic portions in the candidate region, thereby generating a sampled region, and repeating the sampling to generate a plurality of sampled regions;   normalizing a sequence read quantification for each of the sampled regions, thereby generating a normalized quantification for each of the sampled regions;   generating a confidence determination for the candidate region by a process comprising comparing each normalized quantification to a measure of variability of sequence read quantifications for the candidate region for reference samples not having a significant copy number variation in the candidate region; and   providing a classification for presence or absence of the copy number variation for the candidate region for the test sample according to the confidence determination.   
     
     
         10 . The computer-program product of  claim 9 , wherein the actions further comprise:
 sequencing the sample nucleic acid by a sequencer that generates the sequence reads, wherein the sample nucleic acid comprises circulating cell-free nucleic acid from the test sample of a pregnant female bearing a fetus;   mapping, by the computing device, the sequence reads to genomic portions of a reference genome; and   counting, by the computing device, the sequence reads mapped to the genomic portions, wherein the counting generates the quantification of the sequence reads mapped to the genomic portions of the reference genome.   
     
     
         11 . The computer-program product of  claim 9 , wherein the shifting further comprises:
 determining a median read count of genomic portions in the candidate region, thereby providing a candidate region median count;   determining a median read count of genomic portions outside the candidate region and in the selected region, thereby providing an outside median count;   subtracting the outside median count from the candidate region median count, thereby determining a shift.   
     
     
         12 . The computer-program product of  claim 9 , wherein each of the sequence read count quantifications for the candidate region for each of the reference samples is a read count fraction for each candidate region in each of the reference samples. 
     
     
         13 . The computer-program product of  claim 9 , wherein the quantification of sequence reads for genomic portions in the candidate region, outside region and/or selected region is a normalized quantification generated by a normalization process that normalizes GC bias or other bias. 
     
     
         14 . The computer-program product of  claim 9 , wherein the generating the confidence determination further comprises comparing each normalized quantification for each sampled region, or absolute value of each normalized quantification for each sampled region, to a measure of variability threshold for the reference samples. 
     
     
         15 . The computer-program product of  claim 9 , wherein the generating the confidence determination further comprises determining a proportion of the sampled regions for which the normalized quantification is greater than or less than the threshold. 
     
     
         16 . The computer-program product of  claim 15 , wherein the proportion is a ratio of (i) the number of sampled regions for which the normalized quantification is greater than or less than the threshold to (ii) the total number of sampled regions. 
     
     
         17 . A system comprising:
 one or more data processors; and   a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform actions comprising:   selecting a set of genomic portions each coupled to a sequence read quantification for a test sample, wherein:
 the set of genomic portions comprise portions of a reference genome to which sequence reads obtained for sample nucleic acid from the test subject have been mapped, 
 the set of genomic portions comprise (i) genomic portions within a candidate region identified as a candidate copy number variation region, and/or (ii) genomic portions outside the candidate region and within a selected region larger than the candidate region that includes the candidate region, and 
 the candidate region is a chromosome or sub-chromosome region; 
   shifting a quantification of the sequence reads mapped to the genomic portions for the candidate region at or near a quantification of sequence reads mapped to genomic portions outside the candidate region and within the selected region;   sampling the genomic portions from the selected region, wherein a number of sampled genomic portions is about equal to a number of genomic portions in the candidate region, thereby generating a sampled region, and repeating the sampling to generate a plurality of sampled regions;   normalizing a sequence read quantification for each of the sampled regions, thereby generating a normalized quantification for each of the sampled regions;   generating a confidence determination for the candidate region by a process comprising comparing each normalized quantification to a measure of variability of sequence read quantifications for the candidate region for reference samples not having a significant copy number variation in the candidate region; and   
       providing a classification for presence or absence of the copy number variation for the candidate region for the test sample according to the confidence determination. 
     
     
         18 . The system of  claim 17 , wherein the actions further comprise:
 sequencing the sample nucleic acid by a sequencer that generates the sequence reads, wherein the sample nucleic acid comprises circulating cell-free nucleic acid from the test sample of a pregnant female bearing a fetus;   mapping, by the computing device, the sequence reads to genomic portions of a reference genome; and   counting, by the computing device, the sequence reads mapped to the genomic portions, wherein the counting generates the quantification of the sequence reads mapped to the genomic portions of the reference genome.   
     
     
         19 . The system of  claim 17 , wherein the shifting further comprises:
 determining a median read count of genomic portions in the candidate region, thereby providing a candidate region median count;   determining a median read count of genomic portions outside the candidate region and in the selected region, thereby providing an outside median count;   subtracting the outside median count from the candidate region median count, thereby determining a shift.   
     
     
         20 . The system of  claim 17 , wherein each of the sequence read count quantifications for the candidate region for each of the reference samples is a read count fraction for each candidate region in each of the reference samples.

Join the waitlist — get patent alerts

Track US2022093207A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.