US2021272650A1PendingUtilityA1
Methods and Processes for Non-Invasive Assessment of Genetic Variations
Est. expiryOct 10, 2034(~8.2 yrs left)· nominal 20-yr term from priority
G16B 30/10C12Q 1/6869G16B 30/00C12Q 2535/122
74
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Provided herein are methods, processes and apparatuses for non-invasive assessment of genetic variations that make use of nucleic acid fragments from circulating cell free nucleic acid. Also provided herein are methods for partitioning one or more genomic regions of a reference genome into a plurality of portions according to one or more features.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for identifying a presence or absence of a genetic variation, comprising:
a) determining sequencing coverage variability across a reference genome; b) selecting an initial portion length; c) partitioning at least two genomic regions according to the initial portion length in (b); d) comparing the sequencing coverage variability determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of portions for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized portion length; f) re-partitioning at least one of the genomic regions into a plurality of portions according to the optimized portion length in (e), thereby generating at least one re-partitioned genomic region; g) mapping nucleotide sequence reads from a test sample to the plurality of portions of the at least one re-partitioned genomic region, thereby generating mapped nucleotide sequence reads; and h) determining the presence or absence of the genetic variation for the test sample according to raw counts or normalized counts of the mapped nucleotide sequence reads.
2 . The method of claim 1 , wherein determining the sequencing coverage variability in (a) comprises use of a training set of nucleotide sequence reads mapped to portions of a reference genome, which sequence reads are reads of circulating cell-free nucleic acid from a plurality of samples from pregnant females bearing a fetus.
3 . The method of claim 2 , wherein the initial portion length in (b) is selected according to sequencing depth for the training set.
4 . The method of claim 2 , wherein the initial portion length in (b) is selected according to an average fetal fraction for the training set.
5 . The method of claim 4 , wherein the average fetal fraction is determined using the training set.
6 . The method of claim 1 , wherein the initial portion length is between about 1 kb to about 1000 kb.
7 . The method of claim 1 , wherein a total number of portions for a genome is determined according to the initial portion length in (b).
8 . The method of claim 1 , wherein the at least two genomic regions comprise a first genomic region and a second genomic region, and wherein the first genomic region and the second genomic region are substantially similar in size.
9 . The method of claim 8 , wherein the sequencing coverage variability of the first genomic region is determined from a nucleotide sequence read count, or a derivative thereof, for the first genomic region, and wherein sequencing coverage variability of the second genomic region is determined from a nucleotide sequence read count, or a derivative thereof, for the second genomic region.
10 . The method of claim 8 , wherein the sequencing coverage variability of the first genomic region is determined from an average nucleotide sequence read count, or a derivative thereof, for the first genomic region, and wherein sequencing coverage variability of the second genomic region is determined from an average nucleotide sequence read count, or a derivative thereof, for the second genomic region.
11 . The method of claim 1 , further comprising:
i) determining a region-specific fetal fraction for each genomic region according to a correlation between nucleotide sequence read counts per portion and a weighting factor; j) determining a local minimum genomic region size based on the region-specific fetal fraction; and k) adjusting the number of portions for each of the at least two genomic regions to comprise at least two portions based on the local minimum genomic region size, thereby generating a refined re-partitioned genomic region, wherein mapping the nucleotide sequence reads comprises mapping the nucleotide sequence reads from the test sample to the plurality of portions of the at least one refined re-partitioned genomic region from the test sample to the plurality of portions of the at least one re-partitioned genomic region.
12 . The method of claim 1 , further comprising:
i) estimating a fetal fraction for the test sample from a pregnant female bearing a fetus; j) determining a region-specific fetal fraction for each genomic region according to a correlation between nucleotide sequence read counts per portion and a weighting factor; k) determining a local minimum genomic region size based on the region-specific fetal fraction; and l) adjusting the number of portions for each of the at least two genomic regions to comprise at least two portions based on the local minimum genomic region size, thereby generating a refined re-partitioned genomic region, wherein mapping the nucleotide sequence reads comprises mapping the nucleotide sequence reads from the test sample to the plurality of portions of the at least one refined re-partitioned genomic region from the test sample to the plurality of portions of the at least one re-partitioned genomic region.
13 . A system for identifying a presence or absence of a genetic variation, the system comprising:
one or more data processors; and a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform actions including: a) determining sequencing coverage variability across a reference genome; b) selecting an initial portion length; c) partitioning at least two genomic regions according to the initial portion length in (b); d) comparing the sequencing coverage variability determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of portions for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized portion length; f) re-partitioning at least one of the genomic regions into a plurality of portions according to the optimized portion length in (e), thereby generating at least one re-partitioned genomic region; g) mapping nucleotide sequence reads from a test sample to the plurality of portions of the at least one re-partitioned genomic region, thereby generating mapped nucleotide sequence reads; and h) determining the presence or absence of the genetic variation for the test sample according to raw counts or normalized counts of the mapped nucleotide sequence reads.
14 . The system of claim 13 , wherein determining the sequencing coverage variability in (a) comprises use of a training set of nucleotide sequence reads mapped to portions of a reference genome, which sequence reads are reads of circulating cell-free nucleic acid from a plurality of samples from pregnant females bearing a fetus.
15 . The system of claim 14 , wherein the initial portion length in (b) is selected according to sequencing depth for the training set or an average fetal fraction for the training set.
16 . The system of claim 13 , wherein the actions further include:
i) determining a region-specific fetal fraction for each genomic region according to a correlation between nucleotide sequence read counts per portion and a weighting factor; j) determining a local minimum genomic region size based on the region-specific fetal fraction; and k) adjusting the number of portions for each of the at least two genomic regions to comprise at least two portions based on the local minimum genomic region size, thereby generating a refined re-partitioned genomic region, wherein mapping the nucleotide sequence reads comprises mapping the nucleotide sequence reads from the test sample to the plurality of portions of the at least one refined re-partitioned genomic region from the test sample to the plurality of portions of the at least one re-partitioned genomic region.
17 . The system of claim 13 , wherein the actions further include:
i) estimating a fetal fraction for the test sample from a pregnant female bearing a fetus; j) determining a region-specific fetal fraction for each genomic region according to a correlation between nucleotide sequence read counts per portion and a weighting factor; k) determining a local minimum genomic region size based on the region-specific fetal fraction; and l) adjusting the number of portions for each of the at least two genomic regions to comprise at least two portions based on the local minimum genomic region size, thereby generating a refined re-partitioned genomic region, wherein mapping the nucleotide sequence reads comprises mapping the nucleotide sequence reads from the test sample to the plurality of portions of the at least one refined re-partitioned genomic region from the test sample to the plurality of portions of the at least one re-partitioned genomic region.
18 . A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform actions including:
a) determining sequencing coverage variability across a reference genome; b) selecting an initial portion length; c) partitioning at least two genomic regions according to the initial portion length in (b); d) comparing the sequencing coverage variability determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of portions for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized portion length; f) re-partitioning at least one of the genomic regions into a plurality of portions according to the optimized portion length in (e), thereby generating at least one re-partitioned genomic region; g) mapping nucleotide sequence reads from a test sample to the plurality of portions of the at least one re-partitioned genomic region, thereby generating mapped nucleotide sequence reads; and h) determining the presence or absence of the genetic variation for the test sample according to raw counts or normalized counts of the mapped nucleotide sequence reads.
19 . The computer-program product of claim 18 , wherein the actions further include:
i) determining a region-specific fetal fraction for each genomic region according to a correlation between nucleotide sequence read counts per portion and a weighting factor; j) determining a local minimum genomic region size based on the region-specific fetal fraction; and k) adjusting the number of portions for each of the at least two genomic regions to comprise at least two portions based on the local minimum genomic region size, thereby generating a refined re-partitioned genomic region, wherein mapping the nucleotide sequence reads comprises mapping the nucleotide sequence reads from the test sample to the plurality of portions of the at least one refined re-partitioned genomic region from the test sample to the plurality of portions of the at least one re-partitioned genomic region.
20 . The computer-program product of claim 18 , wherein the actions further include:
i) estimating a fetal fraction for the test sample from a pregnant female bearing a fetus; j) determining a region-specific fetal fraction for each genomic region according to a correlation between nucleotide sequence read counts per portion and a weighting factor; k) determining a local minimum genomic region size based on the region-specific fetal fraction; and l) adjusting the number of portions for each of the at least two genomic regions to comprise at least two portions based on the local minimum genomic region size, thereby generating a refined re-partitioned genomic region, wherein mapping the nucleotide sequence reads comprises mapping the nucleotide sequence reads from the test sample to the plurality of portions of the at least one refined re-partitioned genomic region from the test sample to the plurality of portions of the at least one re-partitioned genomic region.Join the waitlist — get patent alerts
Track US2021272650A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.