US2022356527A1PendingUtilityA1

Methods to determine tumor gene copy number by analysis of cell-free dna

Assignee: GUARDANT HEALTH INCPriority: Dec 17, 2015Filed: Dec 16, 2021Published: Nov 10, 2022
Est. expiryDec 17, 2035(~9.4 yrs left)· nominal 20-yr term from priority
C12Q 1/6886C12Q 1/6874G16B 99/00G16B 30/00G16B 20/10G16B 30/10C12Q 1/6809
73
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods are provided herein to improve automatic detection of copy number variation in nucleic acid samples. These methods provide improved approaches for determining baseline copy number of genetic loci within a sample, reduce variation due to features of genetic loci, sample preparation, and probe exhaustion.

Claims

exact text as granted — not AI-modified
1 .- 20 . (canceled) 
     
     
         21 . A method for determining a copy number of one or more genetic loci, comprising:
 (a) receiving, by a computing system having one or more processors and memory, a collection of sequencing reads obtained from a deoxyribonucleic acid (DNA) sample of a subject, wherein the DNA sample is enriched for a plurality of genetic loci using one or more oligonucleotide probes that are complementary to at least a portion of one or more genetic loci from the plurality of genetic loci;   (b) aligning, by the computing system, the collection of sequencing reads to a human reference genome;   (c) generating, by the computing system, from the aligned sequencing reads a first data set of baselining genetic loci comprising, for one or more genetic loci of the plurality of genetic loci, a quantitative measure related to sequencing read coverage of the one or more genetic loci;   (d) adjusting, by the computing system, the first data set of baselining genetic loci into a saturation equilibrium-corrected data set based on a quantitative measure related to guanine-cytosine (GC) content of the one or more genetic loci and a quantitative measure related to a probability that a strand of the DNA molecule derived from the one or more genetic loci of the DNA sample is represented within the sequence reads, wherein the saturation equilibrium-corrected data set comprises a first set of transformed sequencing read coverages of the first data set of genetic loci;   (e) adjusting, by the computing system, the saturation equilibrium-corrected data set into a probe efficiency-corrected data set based on a quantitative measure related to guanine-cytosine (GC) content of the one or more genetic loci of the one or more reference samples and a quantitative measure related to a probability that a strand of DNA molecule derived from the one or more genetic loci of the one or more reference samples is represented within the sequence reads of the one or more reference samples; wherein the probe efficiency-corrected data set comprises a second set of transformed sequencing read coverages of the first data set of genetic loci;   (f) determining, by the computing system, a baseline sequencing read coverage for the first data set, wherein the baseline sequencing read coverage comprises an expected sequencing read coverage for the first data set based on saturation equilibrium and probe efficiency of the one or more oligonucleotide probes that are complementary to the at least the portion of the one or more genetic loci from the plurality of genetic loci; and   (g) applying the baseline sequencing read coverage for the first data set to the probe efficiency-corrected data set to determine a copy number for at least one genetic locus of the one or more genetic loci relative to the baseline sequencing read coverage.   
     
     
         22 . The method of  claim 21 , wherein adjusting the first data set of baselining genetic loci into a saturation equilibrium-corrected data set comprises:
 (i) generating a first transformation by relating the quantitative measure related to the sequencing read coverage in the first data set to both the quantitative measure related to GC content of the one or more genetic loci and the quantitative measure related to a probability that a strand of DNA molecule derived from the one or more genetic loci of the cell-free bodily fluid sample is represented within the sequence reads; and   (ii) applying the first transformation to the sequencing read coverage of the one or more genetic loci of the first data set to generate the saturation equilibrium-corrected data set.   
     
     
         23 . The method of  claim 21 , wherein adjusting the saturation equilibrium-corrected data set into a probe efficiency-corrected data set comprises:
 (i) removing from the saturation-corrected data set genetic loci that are high-variance genetic loci with respect to the first set of transformed sequencing read coverages, thereby providing a second data set of baselining genetic loci;   (ii) generating a reference transformation by relating the sequencing read coverage in the reference data set to both the quantitative measure related to guanine-cytosine (GC) content of the one or more genetic loci GC content of the one or more reference samples and the quantitative measure related to a probability that a strand of DNA molecule derived from the one or more genetic loci of the one or more reference samples is represented within the sequence reads;   (iii) applying the reference transformation to the sequencing read coverage of the one or more genetic loci of the first data set to generate a second transformation; and   (iv) applying the second transformation to the second data set of baselining genetic loci to generate the probe efficiency-corrected data set.   
     
     
         24 . The method of  claim 21 , further comprising, prior to (d), removing genetic loci that are high-variance genetic loci from the first data set, wherein the removing comprises:
 (i) fitting a model relating the quantitative measures related to GC content and the quantitative measures related to sequencing read coverage of the genetic loci; and   (ii) removing from the first data set a subset of the plurality of genetic loci, wherein removing the subset comprises removing at least 5% of the plurality of genetic loci that most differ from the model, thereby providing the first data set of baselining genetic loci.   
     
     
         25 . The method of  claim 23 , wherein removing from the saturation-corrected data set genetic loci that are high-variance genetic loci comprises:
 (i) fitting a model relating the GC content and the first set of transformed sequencing read coverages of the saturation-corrected data set; and   (ii) removing from the saturation-corrected data set a subset of the genetic loci, wherein removing the subset comprises removing at least 5% of the genetic loci that most differ from the model, thereby providing the second data set of baselining genetic loci.   
     
     
         26 . The method of  claim 21 , wherein determining the quantitative measure related to the probability that a strand of DNA derived from the genetic locus is represented within the sequencing reads comprises sorting the sequencing reads into paired reads and unpaired reads, wherein (i) each paired read of the paired reads corresponds to sequencing reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded DNA molecule of the DNA molecules, and (ii) each unpaired read of the unpaired reads represents a first tagged strand having no second differently tagged complementary strand derived from a double-stranded DNA molecule represented among said sequencing reads of the sequencing reads. 
     
     
         27 . The method of  claim 26 , wherein quantitative measures of (i) the paired reads and (ii) the unpaired reads that map to each of one or more genetic loci are determined, to produce a quantitative measure related to total double-stranded DNA molecules of the cell-free bodily fluid sample that map to each of the one or more genetic loci based on the quantitative measures of the paired reads and the unpaired reads mapping to each genetic locus of the one or more genetic loci. 
     
     
         28 . The method of  claim 21 , wherein the sequencing read coverage of the one or more genetic loci is a measure related to central tendency of the sequencing_read coverage of regions of the one or more genetic loci corresponding to the one or more oligonucleotide probes. 
     
     
         29 . The method of  claim 21 , wherein the quantitative measure related to the sequencing read coverage comprises unique molecule counts (UMCs) at one or more genetic loci of the plurality of genetic loci. 
     
     
         30 . The method of  claim 29 , wherein the quantitative measure related to the sequencing read coverage comprises normalizing UMCs by a metric related to library size to provide normalized UMCs. 
     
     
         31 . The method of  claim 30 , wherein normalizing the UMC comprises dividing the UMC of a genetic locus by the sum of all UMCs or dividing the UMC of a genetic locus by the sum of all autosomal UMCs. 
     
     
         32 . The method of  claim 30 , wherein the normalized UMCs further normalized and the further normalization comprises: (i) normalized UMCs are determined for corresponding genetic loci from sequencing reads derived from training samples; (ii) for each genetic locus, normalized UMCs of the sample are further normalized by the median of the normalized UMCs of the training samples at the corresponding loci, thereby providing Relative Abundances (RAs) of genetic loci. 
     
     
         33 . The method of  claim 32 , wherein the quantitative measure related to the sequencing read coverage of the one or more genetic loci comprises the median of the RAs of the probes. 
     
     
         34 . The method of  claim 21 , wherein the quantitative measures of sequencing read coverage are determined for specific regions of a genome. 
     
     
         35 . The method of  claim 33 , wherein the regions can be bins, genes of interest, exons, regions corresponding to the oligonucleotide probes, regions corresponding to primer amplification products, or regions corresponding to primer binding sites. 
     
     
         36 . The method of  claim 33 , wherein the regions of the genome are regions corresponding to the oligonucleotide probes. 
     
     
         37 . The method of  claim 21 , wherein the quantitative measure of the sequencing read coverage comprises a measure related to the count of sequencing reads at one or more genetic loci of the plurality of genetic loci. 
     
     
         38 . The method of  claim 21 , wherein the GC content of one or more genetic loci of the plurality of genetic loci is a measure related to central tendency of GC content of the one or more oligonucleotide probes. 
     
     
         39 . The method of  claim 21 , wherein the DNA sample is a genomic DNA sample obtained from a tissue of the subject. 
     
     
         40 . The method of  claim 21 , wherein the DNA sample is a cell-free DNA sample obtained from a bodily fluid of the subject, wherein the bodily fluid is selected from the group consisting of serum, plasma, urine, and cerebrospinal fluid.

Join the waitlist — get patent alerts

Track US2022356527A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.