US2021358572A1PendingUtilityA1

Methods, systems, and computer-readable media for calculating corrected amplicon coverages

Assignee: LIFE TECHNOLOGIES CORPPriority: Oct 10, 2014Filed: Jul 28, 2021Published: Nov 18, 2021
Est. expiryOct 10, 2034(~8.2 yrs left)· nominal 20-yr term from priority
G16B 30/10G16B 30/20G16B 40/10G16B 30/00G16B 20/20G16B 20/00C12Q 1/6869G16B 40/00
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, systems, and computer-readable media are disclosed for calculating corrected amplicon coverages. One method includes: mapping a plurality of reads of a plurality of amplicons based on amplified target regions of a sample suspected of having one or more genetic abnormalities to a reference sequence that includes one or more nucleic acid sequences corresponding to the amplified target regions; calculating amplicon coverages and total reads, wherein amplicon coverages is a number of reads mapped to an amplicon, and total reads is a number of mapped reads; and calculating corrected amplicon coverages based on the calculated amplicon coverages and calculated total reads by applying a batch effect correction.

Claims

exact text as granted — not AI-modified
1 . A system for identifying a copy number variation, the system including:
 a data storage device that stores instructions; and   a processor configured to execute the instructions, which, when executed by the processor, cause the system to perform a method including:   obtaining, for each training sample of a plurality of training samples in an NGS assay targeting a plurality of amplicons, a plurality of training reads, wherein the plurality of training samples include normal samples having known ploidy, wherein some of the plurality of training samples are prepared in different batches of a plurality of batches than are others of the plurality of training samples;   mapping, for each training sample of the plurality of training samples, the plurality of training reads to a nucleic acid reference sequence corresponding to amplicons of the training sample;   calculating, for each training sample of the plurality of training samples, amplicon coverages and total reads for the training sample, wherein an amplicon coverage is a number of reads mapped to an amplicon and the total reads is a number of mapped reads;   representing amplicon coverages for each training sample of the plurality of training samples as a vector to obtain a plurality of vectors representing amplicon coverages of the plurality of training samples in the plurality of batches;   determining values of a batch effect by applying a principal components analysis to the plurality of vectors, wherein each principal component is used to obtain a vector of batch effect values;   obtaining, from each test sample of a plurality of test samples in the NGS assay targeting the plurality of amplicons, a plurality of reads of a plurality of amplicons based on amplified target regions of the test sample;   mapping, for each test sample of the plurality of test samples, the plurality of reads to a reference sequence, the reference sequence including one or more nucleic acid sequences corresponding to the amplified target regions;   calculating, for each test sample of the plurality of test samples, amplicon coverages and total reads for the test sample;   representing amplicon coverages for each test sample of the plurality of test samples as a test vector to obtain a plurality of test vectors representing amplicon coverages of the plurality of test samples;   projecting each of the plurality of test vectors onto the principal components identified by the principal components analysis of the training samples to determine a plurality of scaling factors;   calculating corrected amplicon coverages by applying a batch effect correction for each test sample based on the calculated amplicon coverages for the test sample, the calculated total reads for the test sample, the scaling factor determined for the test sample and the batch effect values corresponding to the principal components determined for the training sample; and   identifying the copy number variation of the test sample based on a likelihood of a ploidy state for the corrected amplicon coverages.   
     
     
         2 . The system of  claim 1 , wherein the processor is further configured to execute the instructions to perform the method including:
 amplifying the target regions of nucleic acids isolated from the test sample to produce the plurality of amplicons based on the amplified target regions of the test sample; and   sequencing the plurality of amplicons to obtain the plurality of reads.   
     
     
         3 . The system of  claim 2 , wherein amplifying the target regions of nucleic acids isolated from the test sample includes multiplex amplification. 
     
     
         4 . The system of  claim 1 , wherein the processor is further configured to execute the instructions to perform the method including:
 determining a maximum score path for the plurality of amplicons based on the likelihoods calculated for a range of ploidy states; and   identifying the copy number variations based on the maximum score path.   
     
     
         5 . The system of  claim 4 , wherein the processor is further configured to execute the instructions to perform the method including:
 normalizing, prior to calculating the likelihoods for the maximum score path, the corrected amplicon coverages based on the total reads for the test sample.   
     
     
         6 . The system of  claim 1 , wherein the step of calculating corrected amplicon coverages further includes determining a logarithm of the corrected copy number for an i-th amplicon based on a product of the scaling factor and a logarithm of the batch effect value determined for the i-th amplicon. 
     
     
         7 . The system of  claim 1 , wherein the step of representing amplicon coverages for each training sample, a vector element represents a difference in the amplicon coverages for a pair of adjacent amplicons of the training sample. 
     
     
         8 . The system of  claim 1 , wherein the step of representing amplicon coverages for each test sample, a test vector element represents a difference in the amplicon coverages for a pair of adjacent amplicons of the test sample 
     
     
         9 . A non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for identifying a copy number variation, the method including:
 obtaining, for each training sample of a plurality of training samples in an NGS assay targeting a plurality of amplicons, a plurality of training reads, wherein the plurality of training samples include normal samples having known ploidy, wherein some of the plurality of training samples are prepared in different batches of a plurality of batches than are others of the plurality of training samples;   mapping, for each training sample of the plurality of training samples, the plurality of training reads to a nucleic acid reference sequence corresponding to amplicons of the training sample;   calculating, for each training sample of the plurality of training samples, amplicon coverages and total reads for the training sample, wherein an amplicon coverage is a number of reads mapped to an amplicon and the total reads is a number of mapped reads;   representing amplicon coverages for each training sample of the plurality of training samples as a vector to obtain a plurality of vectors representing amplicon coverages of the plurality of training samples in the plurality of batches;   determining values of a batch effect by applying a principal components analysis to the plurality of vectors, wherein each principal component is used to obtain a vector of batch effect values;   obtaining, from each test sample of a plurality of test samples in the NGS assay targeting the plurality of amplicons, a plurality of reads of a plurality of amplicons based on amplified target regions of the test sample;   mapping, for each test sample of the plurality of test samples, the plurality of reads to a reference sequence, the reference sequence including one or more nucleic acid sequences corresponding to the amplified target regions;   calculating, for each test sample of the plurality of test samples, amplicon coverages and total reads for the test sample;   representing amplicon coverages for each test sample of the plurality of test samples as a test vector to obtain a plurality of test vectors representing amplicon coverages of the plurality of test samples;   projecting each of the plurality of test vectors onto the principal components identified by the principal components analysis of the training samples to determine a plurality of scaling factors;   calculating corrected amplicon coverages by applying a batch effect correction for each test sample based on the calculated amplicon coverages for the test sample, the calculated total reads for the test sample, the scaling factor determined for the test sample and the batch effect values corresponding to the principal components determined for the training sample; and   identifying the copy number variation of the test sample based on a likelihood of a ploidy state for the corrected amplicon coverages.   
     
     
         10 . The computer-readable medium of  claim 9 , further comprising instructions for:
 amplifying the target regions of nucleic acids isolated from the test sample to produce the plurality of amplicons based on the amplified target regions of the test sample; and   sequencing the plurality of amplicons to obtain the plurality of reads.   
     
     
         11 . The computer-readable medium of  claim 10 , wherein amplifying the target regions of nucleic acids isolated from the test sample includes multiplex amplification. 
     
     
         12 . The computer-readable medium of  claim 9 , further comprising instructions for:
 determining a maximum score path for the plurality of amplicons based on the likelihoods calculated for a range of ploidy states; and   identifying the copy number variations based on the maximum score path.   
     
     
         13 . The computer-readable medium of  claim 12 , further comprising instructions for:
 normalizing, prior to calculating the likelihoods for the maximum score path, the corrected amplicon coverages based on the total reads for the test sample.   
     
     
         14 . The computer-readable medium of  claim 9 , wherein the step of calculating corrected amplicon coverages further includes determining a logarithm of the corrected copy number for an i-th amplicon based on a product of the scaling factor and a logarithm of the batch effect value determined for the i-th amplicon 
     
     
         15 . The computer-readable medium of  claim 9 , wherein the step of representing amplicon coverages for each training sample, a vector element represents a difference in the amplicon coverages for a pair of adjacent amplicons of the training sample. 
     
     
         16 . The system of  claim 9 , wherein for step of representing amplicon coverages for each test sample, a test vector element represents a difference in the amplicon coverages for a pair of adjacent amplicons of the test sample. 
     
     
         17 . A method for for identifying a copy number variation, comprising
 mapping a plurality of reads of amplicons of a test sample obtained using an NGS assay to a reference sequence, the reference sequence including one or more nucleic acid sequences corresponding to amplified target regions of the test sample, wherein the test sample has unknown ploidy;   calculating amplicon coverages and total reads for the test sample;   representing the amplicon coverages for the test sample as a test vector ;   projecting the test vector onto principal components to obtain a scaling factor, wherein the principal components correspond to batch effect values determined by a principal components analysis of a plurality of training reads obtained from a plurality of training samples using the NGS assay, wherein the plurality of training samples have known ploidy;   calculating corrected amplicon coverages by applying a batch effect correction for the test sample based on the calculated amplicon coverages for the test sample, the calculated total reads for the test sample, the scaling factor determined for the test sample and the batch effect values corresponding to the principal components determined for the training samples; and   identifying the copy number variation of the test sample based on a likelihood of a ploidy state for the corrected amplicon coverages.   
     
     
         18 . The method of  claim 17 , further comprising:
 amplifying the target regions of nucleic acids isolated from the test sample to produce the plurality of amplicons based on the amplified target regions of the test sample; and   sequencing the plurality of amplicons to obtain the plurality of reads.   
     
     
         19 . The method of  claim 17 , wherein the step of identifying the copy number variation further comprises calculating a likelihood of a non-integer ploidy state. 
     
     
         20 . The method of  claim 17 , wherein the step of identifying the copy number variation further comprises calculating the likelihood for a plurality of ploidy states within a range of ploidy states.

Join the waitlist — get patent alerts

Track US2021358572A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.