Methods, systems, and computer-readable media for calculating corrected amplicon coverages
Abstract
Methods, systems, and computer-readable media are disclosed for calculating corrected amplicon coverages. One method includes: mapping a plurality of reads of a plurality of amplicons based on amplified target regions of a sample suspected of having one or more genetic abnormalities to a reference sequence that includes one or more nucleic acid sequences corresponding to the amplified target regions; calculating amplicon coverages and total reads, wherein amplicon coverages is a number of reads mapped to an amplicon, and total reads is a number of mapped reads; and calculating corrected amplicon coverages based on the calculated amplicon coverages and calculated total reads by applying a batch effect correction.
Claims
exact text as granted — not AI-modified1 . A system for identifying a copy number variation, the system including:
a data storage device that stores instructions; and a processor configured to execute the instructions, which, when executed by the processor, cause the system to perform a method including: obtaining, for each training sample of a plurality of training samples in an NGS assay targeting a plurality of amplicons, a plurality of training reads, wherein the plurality of training samples include normal samples having known ploidy, wherein some of the plurality of training samples are prepared in different batches of a plurality of batches than are others of the plurality of training samples; mapping, for each training sample of the plurality of training samples, the plurality of training reads to a nucleic acid reference sequence corresponding to amplicons of the training sample; calculating, for each training sample of the plurality of training samples, amplicon coverages and total reads for the training sample, wherein an amplicon coverage is a number of reads mapped to an amplicon and the total reads is a number of mapped reads; representing amplicon coverages for each training sample of the plurality of training samples as a vector to obtain a plurality of vectors representing amplicon coverages of the plurality of training samples in the plurality of batches; determining values of a batch effect by applying a principal components analysis to the plurality of vectors, wherein each principal component is used to obtain a vector of batch effect values; obtaining, from each test sample of a plurality of test samples in the NGS assay targeting the plurality of amplicons, a plurality of reads of a plurality of amplicons based on amplified target regions of the test sample; mapping, for each test sample of the plurality of test samples, the plurality of reads to a reference sequence, the reference sequence including one or more nucleic acid sequences corresponding to the amplified target regions; calculating, for each test sample of the plurality of test samples, amplicon coverages and total reads for the test sample; representing amplicon coverages for each test sample of the plurality of test samples as a test vector to obtain a plurality of test vectors representing amplicon coverages of the plurality of test samples; projecting each of the plurality of test vectors onto the principal components identified by the principal components analysis of the training samples to determine a plurality of scaling factors; calculating corrected amplicon coverages by applying a batch effect correction for each test sample based on the calculated amplicon coverages for the test sample, the calculated total reads for the test sample, the scaling factor determined for the test sample and the batch effect values corresponding to the principal components determined for the training sample; and identifying the copy number variation of the test sample based on a likelihood of a ploidy state for the corrected amplicon coverages.
2 . The system of claim 1 , wherein the processor is further configured to execute the instructions to perform the method including:
amplifying the target regions of nucleic acids isolated from the test sample to produce the plurality of amplicons based on the amplified target regions of the test sample; and sequencing the plurality of amplicons to obtain the plurality of reads.
3 . The system of claim 2 , wherein amplifying the target regions of nucleic acids isolated from the test sample includes multiplex amplification.
4 . The system of claim 1 , wherein the processor is further configured to execute the instructions to perform the method including:
determining a maximum score path for the plurality of amplicons based on the likelihoods calculated for a range of ploidy states; and identifying the copy number variations based on the maximum score path.
5 . The system of claim 4 , wherein the processor is further configured to execute the instructions to perform the method including:
normalizing, prior to calculating the likelihoods for the maximum score path, the corrected amplicon coverages based on the total reads for the test sample.
6 . The system of claim 1 , wherein the step of calculating corrected amplicon coverages further includes determining a logarithm of the corrected copy number for an i-th amplicon based on a product of the scaling factor and a logarithm of the batch effect value determined for the i-th amplicon.
7 . The system of claim 1 , wherein the step of representing amplicon coverages for each training sample, a vector element represents a difference in the amplicon coverages for a pair of adjacent amplicons of the training sample.
8 . The system of claim 1 , wherein the step of representing amplicon coverages for each test sample, a test vector element represents a difference in the amplicon coverages for a pair of adjacent amplicons of the test sample
9 . A non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for identifying a copy number variation, the method including:
obtaining, for each training sample of a plurality of training samples in an NGS assay targeting a plurality of amplicons, a plurality of training reads, wherein the plurality of training samples include normal samples having known ploidy, wherein some of the plurality of training samples are prepared in different batches of a plurality of batches than are others of the plurality of training samples; mapping, for each training sample of the plurality of training samples, the plurality of training reads to a nucleic acid reference sequence corresponding to amplicons of the training sample; calculating, for each training sample of the plurality of training samples, amplicon coverages and total reads for the training sample, wherein an amplicon coverage is a number of reads mapped to an amplicon and the total reads is a number of mapped reads; representing amplicon coverages for each training sample of the plurality of training samples as a vector to obtain a plurality of vectors representing amplicon coverages of the plurality of training samples in the plurality of batches; determining values of a batch effect by applying a principal components analysis to the plurality of vectors, wherein each principal component is used to obtain a vector of batch effect values; obtaining, from each test sample of a plurality of test samples in the NGS assay targeting the plurality of amplicons, a plurality of reads of a plurality of amplicons based on amplified target regions of the test sample; mapping, for each test sample of the plurality of test samples, the plurality of reads to a reference sequence, the reference sequence including one or more nucleic acid sequences corresponding to the amplified target regions; calculating, for each test sample of the plurality of test samples, amplicon coverages and total reads for the test sample; representing amplicon coverages for each test sample of the plurality of test samples as a test vector to obtain a plurality of test vectors representing amplicon coverages of the plurality of test samples; projecting each of the plurality of test vectors onto the principal components identified by the principal components analysis of the training samples to determine a plurality of scaling factors; calculating corrected amplicon coverages by applying a batch effect correction for each test sample based on the calculated amplicon coverages for the test sample, the calculated total reads for the test sample, the scaling factor determined for the test sample and the batch effect values corresponding to the principal components determined for the training sample; and identifying the copy number variation of the test sample based on a likelihood of a ploidy state for the corrected amplicon coverages.
10 . The computer-readable medium of claim 9 , further comprising instructions for:
amplifying the target regions of nucleic acids isolated from the test sample to produce the plurality of amplicons based on the amplified target regions of the test sample; and sequencing the plurality of amplicons to obtain the plurality of reads.
11 . The computer-readable medium of claim 10 , wherein amplifying the target regions of nucleic acids isolated from the test sample includes multiplex amplification.
12 . The computer-readable medium of claim 9 , further comprising instructions for:
determining a maximum score path for the plurality of amplicons based on the likelihoods calculated for a range of ploidy states; and identifying the copy number variations based on the maximum score path.
13 . The computer-readable medium of claim 12 , further comprising instructions for:
normalizing, prior to calculating the likelihoods for the maximum score path, the corrected amplicon coverages based on the total reads for the test sample.
14 . The computer-readable medium of claim 9 , wherein the step of calculating corrected amplicon coverages further includes determining a logarithm of the corrected copy number for an i-th amplicon based on a product of the scaling factor and a logarithm of the batch effect value determined for the i-th amplicon
15 . The computer-readable medium of claim 9 , wherein the step of representing amplicon coverages for each training sample, a vector element represents a difference in the amplicon coverages for a pair of adjacent amplicons of the training sample.
16 . The system of claim 9 , wherein for step of representing amplicon coverages for each test sample, a test vector element represents a difference in the amplicon coverages for a pair of adjacent amplicons of the test sample.
17 . A method for for identifying a copy number variation, comprising
mapping a plurality of reads of amplicons of a test sample obtained using an NGS assay to a reference sequence, the reference sequence including one or more nucleic acid sequences corresponding to amplified target regions of the test sample, wherein the test sample has unknown ploidy; calculating amplicon coverages and total reads for the test sample; representing the amplicon coverages for the test sample as a test vector ; projecting the test vector onto principal components to obtain a scaling factor, wherein the principal components correspond to batch effect values determined by a principal components analysis of a plurality of training reads obtained from a plurality of training samples using the NGS assay, wherein the plurality of training samples have known ploidy; calculating corrected amplicon coverages by applying a batch effect correction for the test sample based on the calculated amplicon coverages for the test sample, the calculated total reads for the test sample, the scaling factor determined for the test sample and the batch effect values corresponding to the principal components determined for the training samples; and identifying the copy number variation of the test sample based on a likelihood of a ploidy state for the corrected amplicon coverages.
18 . The method of claim 17 , further comprising:
amplifying the target regions of nucleic acids isolated from the test sample to produce the plurality of amplicons based on the amplified target regions of the test sample; and sequencing the plurality of amplicons to obtain the plurality of reads.
19 . The method of claim 17 , wherein the step of identifying the copy number variation further comprises calculating a likelihood of a non-integer ploidy state.
20 . The method of claim 17 , wherein the step of identifying the copy number variation further comprises calculating the likelihood for a plurality of ploidy states within a range of ploidy states.Join the waitlist — get patent alerts
Track US2021358572A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.