US2014229117A1PendingUtilityA1

Methods for estimating genome-wide copy number variations

Assignee: COMPLETE GENOMICS INCPriority: Oct 13, 2010Filed: Apr 15, 2014Published: Aug 14, 2014
Est. expiryOct 13, 2030(~4.2 yrs left)· nominal 20-yr term from priority
G16B 25/20G16B 30/10G16B 20/10G16B 30/20G16B 20/20G16B 40/30G16B 30/00G16B 40/00G16B 25/00C12Q 1/6809G16B 20/00C12Q 1/6827G06F 19/18
66
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods for determining the copy number of a genomic region at a detection position of a target sequence in a sample are disclosed. Genomic regions of a target sequence in a sample are sequenced and measurement data for sequence coverage is obtained. Sequence coverage bias is corrected and may be normalized against a baseline sample. Hidden Markov Model (HMM) segmentation, scoring, and output are performed, and in some embodiments population-based no-calling and identification of low-confidence regions may also be performed. A total copy number value and region-specific copy number value for a plurality of regions are then estimated.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A non-transitory computer-readable medium comprising instructions tangibly embodied thereon, the instructions when executed by a computer processor causing the processor to perform the operations of:
 obtaining, using the computer processor, coverage values for each given position in a baseline or reference sample for the sequence coverage of said target polynucleotide using data generated from mate-pair mappings;   correcting, using the computer processor, the coverage values for each given position for sequence coverage bias, wherein correcting the coverage values for each given position in a baseline or reference sample comprises performing ploidy-aware baseline correction; and   estimating, using the computer processor, a total copy number value and region-specific copy number value for each of a plurality of genomic regions based at least on the corrected coverage values for each given position in a baseline or reference sample.   
     
     
         2 . A non-transitory computer-readable medium comprising instructions tangibly embodied thereon, the instructions when executed by a computer processor causing the processor to perform the operations of:
 obtaining, using the computer processor, coverage values for each given position in a baseline or reference sample for the sequence coverage of said target polynucleotide using data generated from mate-pair mappings;   correcting, using the computer processor, the coverage values for each given position for sequence coverage bias, wherein correcting the coverage values for each given position in a baseline or reference sample comprises performing ploidy-aware baseline correction; and   performing Hidden Markov Model (HMM) segmentation, scoring, and output based on the corrected coverage values for each given position in a baseline or reference sample;   based on the HMM scoring and output, performing population-based no-calling and identification of low-confidence regions; and   based on the HMM scoring and output, estimating a total copy number value and region-specific copy number value for a plurality of regions.   
     
     
         3 . A system of determining copy number variation of a genomic region at a detection position of a target polynucleotide sequence, comprising:
 a computer processor; and   a computer-readable storage medium coupled to said computer processor, the storage medium having instructions tangibly embodied thereon, the instructions when executed by said processor causing said processor to perform the operations of:   obtaining, using the computer processor, coverage values for each given position in a baseline or reference sample for the sequence coverage of said target polynucleotide using data generated from mate-pair mappings;   correcting, using the computer processor, the coverage values for each given position for sequence coverage bias, wherein correcting the coverage values for each given position in a baseline or reference sample comprises performing ploidy-aware baseline correction; and   estimating, using a computer, a total copy number value and region-specific copy number value for each of a plurality of genomic regions based at least on the corrected coverage values for each given position in a baseline or reference sample.

Join the waitlist — get patent alerts

Track US2014229117A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.