US2020335178A1PendingUtilityA1

Detecting repeat expansions with short read sequencing data

Assignee: ILLUMINA CAMBRIDGE LTDPriority: Sep 12, 2014Filed: May 7, 2020Published: Oct 22, 2020
Est. expirySep 12, 2034(~8.1 yrs left)· nominal 20-yr term from priority
G16B 30/00G16B 30/10G16B 20/20G16B 20/00C12Q 2600/156C12Q 1/6883
60
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The disclosed embodiments concern methods, apparatus, systems and computer program products for determining the presence or absence of repeat expansions of interest, including repeat expansions of repeat sequences that are medically significant. Some embodiments provide methods for identifying and calling medically relevant repeat expansions using anchored reads. An anchored read is a paired end read that is unaligned to a repeat sequence under consideration, but it is paired with an anchor read that is aligned to or near the repeat sequence. Some embodiments use both anchor and anchored reads to determine the presence or absence of the repeat expansions. System, apparatus, and computer program products are also provided for determining repeat expansion implementing the methods disclosed.

Claims

exact text as granted — not AI-modified
1 . A method, implemented using a computer system comprising one or more processors and system memory, for determining the presence or absence of a repeat expansion of a repeat sequence in a test sample comprising nucleic acids, wherein the repeat sequence comprises repeats of a repeat unit of nucleotides, the method comprising:
 (a) obtaining at least 100,000 paired end reads of the test sample, wherein the paired end reads have been processed to align to a reference sequence comprising the repeat sequence;   (b) comparing, by the one or more processors, genomic locations of the at least 100,000 paired end reads to the genomic location of the repeat sequence to identify anchor and anchored reads in the paired end reads, wherein the anchor reads are aligned reads having genomic locations that are the same as or near the genomic location of the repeat sequence, and the anchored reads are unaligned reads that are paired with the anchor reads; and   (c) determining if the repeat expansion is likely to be present in the test sample based at least in part on numbers of repeats of the repeat unit in the anchored reads.   
     
     
         2 . The method of  claim 1 , wherein (c) comprises determining if the repeat expansion is likely to be present in the test sample based on the anchor reads and the anchored reads. 
     
     
         3 . (canceled) 
     
     
         4 . The method of  claim 1 , wherein (c) comprises:
 obtaining the number of identified reads that are high-count reads, wherein the high-count reads comprise reads having more repeats than a threshold value; and   comparing the number of high-count reads in the test sample to a call criterion.   
     
     
         5 . The method of  claim 4 , wherein the threshold value for high-count reads is at least about 80% of the maximum number of repeats, which maximum is calculated from the length of the paired end reads and the length of the repeat unit. 
     
     
         6 . The method of  claim 5 , wherein the threshold value for high-count reads is at least about 90% of the maximum number of repeats. 
     
     
         7 . The method of  claim 4 , wherein the call criterion is obtained from a distribution of high-count reads of control samples. 
     
     
         8 . The method of  claim 1 , wherein the anchor reads are aligned to or within about 5 kb of the repeat sequence. 
     
     
         9 . The method of  claim 1 , wherein the anchor reads are aligned to or within about 1 kb of the repeat sequence. 
     
     
         10 . The method of  claim 1 , wherein the unaligned reads comprise reads that cannot be aligned or are poorly aligned to the reference sequence. 
     
     
         11 . The method of  claim 1 , wherein (c) comprises comparing a distribution of numbers of repeats of the repeat unit in the identified reads for the test sample and a distribution of numbers of repeats for one or more control samples. 
     
     
         12 . The method of  claim 11 , wherein comparing the distribution for the test sample to the distribution for the control samples comprises using a Mann-Whitney rank test to determine if the distribution of the test sample statistically significantly differs from the distribution of the control samples. 
     
     
         13 . The method of  claim 12 , further comprising determining that the repeat expansion is likely present in the test sample if the test sample's distribution is skewed more towards higher numbers of repeats than the control samples, and the p value for the Mann-Whitney rank test is smaller than about 0.0001. 
     
     
         14 . The method of  claim 13 , further comprising determining that the repeat expansion is likely present in the test sample if the test sample's distribution is skewed more towards higher numbers of repeats than the control samples, and the p value for the Mann-Whitney rank test is smaller than about 0.00001. 
     
     
         15 . The method of any of the preceding claims, wherein the numbers of repeats are numbers of in-frame repeats. 
     
     
         16 . The method of any of the preceding claims, further comprising using a sequencer to generate paired end reads from the test sample. 
     
     
         17 . The method of any of the preceding claims, further comprising extracting the test sample from an individual. 
     
     
         18 . A method for detecting a repeat expansion in a test sample comprising nucleic acids, the method comprising:
 (a) obtaining at least 100,000 paired end reads of the test sample;   (b) aligning the at least 100,000 paired end reads to a reference genome;   (c) identifying unaligned reads from the whole genome, wherein the unaligned reads comprise paired end reads that cannot be aligned or are poorly aligned to the reference genome; and   (d) analyzing the numbers of repeats of a repeat unit in the unaligned reads to determine if a repeat expansion is likely present in the test sample.   
     
     
         19 . The method of  claim 18 , wherein analyzing the numbers of repeats of the repeat unit in the unaligned reads comprises:
 obtaining the number of high-count reads, wherein the high-count reads comprise unaligned reads having more repeats than a threshold value; and   comparing the number of high-count reads in the test sample to a call criterion.   
     
     
         20 . The method of  claim 19 , wherein the threshold value for high-count reads is at least about 80% of the maximum number of repeats, which maximum is calculated as the ratio of the length of the paired end reads over the length of the repeat unit. 
     
     
         21 . A system for determining the presence or absence of a repeat expansion of a repeat sequence in a test sample comprising nucleic acids, wherein the repeat sequence comprises repeats of a repeat unit, the method comprising:
 a sequencer for sequencing nucleic acids of the test sample;   a processor; and   one or more computer-readable storage media having stored thereon instructions for execution on said processor to evaluate copy number in the test sample by:
 (a) obtaining at least 100,000 paired end reads of the test sample, wherein the paired end reads have been processed to align to a reference sequence comprising the repeat sequence; 
 (b) comparing, by the one or more processors, genomic locations of the at least 100,000 paired end reads to the genomic location of the repeat sequence to identify anchor and anchored reads in the paired end reads, wherein the anchor reads are aligned reads having genomic locations that are the same as or near the genomic location of the repeat sequence, and the anchored reads are unaligned reads that are paired with the anchor reads; and 
 (c) determining if the repeat expansion is likely to be present in the test sample based at least in part on numbers of repeats of the repeat unit in the anchored reads.

Join the waitlist — get patent alerts

Track US2020335178A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.