Detecting repeat expansions with short read sequencing data
Abstract
The disclosed embodiments concern methods, apparatus, systems and computer program products for determining the presence or absence of repeat expansions of interest, including repeat expansions of repeat sequences that are medically significant. Some embodiments provide methods for identifying and calling medically relevant repeat expansions using anchored reads. An anchored read is a paired end read that is unaligned to a repeat sequence under consideration, but it is paired with an anchor read that is aligned to or near the repeat sequence. Some embodiments use both anchor and anchored reads to determine the presence or absence of the repeat expansions. System, apparatus, and computer program products are also provided for determining repeat expansion implementing the methods disclosed.
Claims
exact text as granted — not AI-modified1 . A method, implemented using a computer system comprising one or more processors and system memory, for determining the presence or absence of a repeat expansion of a repeat sequence in a test sample comprising nucleic acids, wherein the repeat sequence comprises repeats of a repeat unit of nucleotides, the method comprising:
(a) obtaining at least 100,000 paired end reads of the test sample, wherein the paired end reads have been processed to align to a reference sequence comprising the repeat sequence; (b) comparing, by the one or more processors, genomic locations of the at least 100,000 paired end reads to the genomic location of the repeat sequence to identify anchor and anchored reads in the paired end reads, wherein the anchor reads are aligned reads having genomic locations that are the same as or near the genomic location of the repeat sequence, and the anchored reads are unaligned reads that are paired with the anchor reads; and (c) determining if the repeat expansion is likely to be present in the test sample based at least in part on numbers of repeats of the repeat unit in the anchored reads.
2 . The method of claim 1 , wherein (c) comprises determining if the repeat expansion is likely to be present in the test sample based on the anchor reads and the anchored reads.
3 . (canceled)
4 . The method of claim 1 , wherein (c) comprises:
obtaining the number of identified reads that are high-count reads, wherein the high-count reads comprise reads having more repeats than a threshold value; and comparing the number of high-count reads in the test sample to a call criterion.
5 . The method of claim 4 , wherein the threshold value for high-count reads is at least about 80% of the maximum number of repeats, which maximum is calculated from the length of the paired end reads and the length of the repeat unit.
6 . The method of claim 5 , wherein the threshold value for high-count reads is at least about 90% of the maximum number of repeats.
7 . The method of claim 4 , wherein the call criterion is obtained from a distribution of high-count reads of control samples.
8 . The method of claim 1 , wherein the anchor reads are aligned to or within about 5 kb of the repeat sequence.
9 . The method of claim 1 , wherein the anchor reads are aligned to or within about 1 kb of the repeat sequence.
10 . The method of claim 1 , wherein the unaligned reads comprise reads that cannot be aligned or are poorly aligned to the reference sequence.
11 . The method of claim 1 , wherein (c) comprises comparing a distribution of numbers of repeats of the repeat unit in the identified reads for the test sample and a distribution of numbers of repeats for one or more control samples.
12 . The method of claim 11 , wherein comparing the distribution for the test sample to the distribution for the control samples comprises using a Mann-Whitney rank test to determine if the distribution of the test sample statistically significantly differs from the distribution of the control samples.
13 . The method of claim 12 , further comprising determining that the repeat expansion is likely present in the test sample if the test sample's distribution is skewed more towards higher numbers of repeats than the control samples, and the p value for the Mann-Whitney rank test is smaller than about 0.0001.
14 . The method of claim 13 , further comprising determining that the repeat expansion is likely present in the test sample if the test sample's distribution is skewed more towards higher numbers of repeats than the control samples, and the p value for the Mann-Whitney rank test is smaller than about 0.00001.
15 . The method of any of the preceding claims, wherein the numbers of repeats are numbers of in-frame repeats.
16 . The method of any of the preceding claims, further comprising using a sequencer to generate paired end reads from the test sample.
17 . The method of any of the preceding claims, further comprising extracting the test sample from an individual.
18 . A method for detecting a repeat expansion in a test sample comprising nucleic acids, the method comprising:
(a) obtaining at least 100,000 paired end reads of the test sample; (b) aligning the at least 100,000 paired end reads to a reference genome; (c) identifying unaligned reads from the whole genome, wherein the unaligned reads comprise paired end reads that cannot be aligned or are poorly aligned to the reference genome; and (d) analyzing the numbers of repeats of a repeat unit in the unaligned reads to determine if a repeat expansion is likely present in the test sample.
19 . The method of claim 18 , wherein analyzing the numbers of repeats of the repeat unit in the unaligned reads comprises:
obtaining the number of high-count reads, wherein the high-count reads comprise unaligned reads having more repeats than a threshold value; and comparing the number of high-count reads in the test sample to a call criterion.
20 . The method of claim 19 , wherein the threshold value for high-count reads is at least about 80% of the maximum number of repeats, which maximum is calculated as the ratio of the length of the paired end reads over the length of the repeat unit.
21 . A system for determining the presence or absence of a repeat expansion of a repeat sequence in a test sample comprising nucleic acids, wherein the repeat sequence comprises repeats of a repeat unit, the method comprising:
a sequencer for sequencing nucleic acids of the test sample; a processor; and one or more computer-readable storage media having stored thereon instructions for execution on said processor to evaluate copy number in the test sample by:
(a) obtaining at least 100,000 paired end reads of the test sample, wherein the paired end reads have been processed to align to a reference sequence comprising the repeat sequence;
(b) comparing, by the one or more processors, genomic locations of the at least 100,000 paired end reads to the genomic location of the repeat sequence to identify anchor and anchored reads in the paired end reads, wherein the anchor reads are aligned reads having genomic locations that are the same as or near the genomic location of the repeat sequence, and the anchored reads are unaligned reads that are paired with the anchor reads; and
(c) determining if the repeat expansion is likely to be present in the test sample based at least in part on numbers of repeats of the repeat unit in the anchored reads.Join the waitlist — get patent alerts
Track US2020335178A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.