Methods for detecting copy-number variations in next-generation sequencing
Abstract
Copy Number Variants (CNV) detection methods may integrate CNV detection into workflow for next generation sequencer (NGS) data processing, in parallel with SNP and INDEL variant calling. CNV detection methods may analyze coverage patterns across a set of genomic regions and across samples from different patients. The methods do not require specifically chosen reference samples as, but automatically select reference samples from the same batch, for each sample tested. CNV detection methods may detect CNVs in a set of samples without assumptions about CNV status of any of those samples. Embodiments herein may apply the CNV detection scheme iteratively to improve detection performance. The proposed methods may further comprise the step of iteratively feeding back information about the CNVs in the samples into the next iteration step. The methods may also use information from the NGS workflow, such as information on SNP fractions, as input to the NGS CNV detection.
Claims
exact text as granted — not AI-modifiedWhat is claimed:
1 . A method for detecting copy-number values (CNV) from a pool of DNA samples enriched with a target enrichment technology, each enriched DNA sample being associated with a library of pooled fragments from a set of amplicons/regions, each amplicon/region being sequenced with a high-throughput sequencer to generate coverage count for each sample and for each amplicon/region, comprising:
normalizing, with a data processing unit, the coverage count associated with each sample; selecting, with a data processing unit, for each sample, a set of reference samples as the samples with the closest normalized coverage count to the normalized coverage count of said sample, the number of reference samples in each subset of reference samples being a function of the total number of samples and being smaller than the total number of samples; for each sample, estimating the copy-number values in said sample as a function of at least the coverage counts in said sample and of at least the coverage counts in the selected set of reference samples for said sample.
2 . The method of claim 1 , wherein the number of reference samples N R in each set of reference samples is given by N R =10.25*M+2, where N is the total number of samples.
3 . The method of any of the preceding claims, further comprising: for each sample and for each amplicon/region, estimating the likelihood for each possible copy-number value.
4 . The method of claim 3 , wherein a Hidden Markov Model is further used to estimate the copy-number values and their confidence levels for each amplicon/region.
5 . The method of claim 4 , further comprising: excluding possible copy number values for which the confidence level is below a minimum threshold.
6 . The method of claim 1 , wherein the estimate of the copy-number values is calculated using information on the SNP fractions and coverage.
7 . The method of claim 1 , further comprising: applying a principal-component filter to the coverage count.
8 . The method of claim 1 , wherein normalizing the coverage count associated with each sample depends on a prior estimate of the copy number values for each sample and each amplicon/region.
9 . The method of claim 1 , wherein selecting a set of reference samples depends on a prior estimate of the copy number values for each sample and each amplicon/region.Join the waitlist — get patent alerts
Track US2022130488A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.