Comprehensive analysis pipeline for discovery of human genetic variation
Abstract
Systems and methods for analyzing genetic sequence data involve: (a) obtaining, by a computer system, genetic sequencing data pertaining to a subject; (b) splitting the genetic sequencing data into a plurality of segments; (c) processing the genetic sequencing data such that intra-segment reads, read pairs with both mates mapped to the same data set, are saved to a respective plurality of individual binary alignment map (BAM) files corresponding to that respective segment; (d) processing the genetic sequencing data such that inter-segment reads, read pairs with both mates mapped to different segments, are saved into at least a second BAM file; and (e) processing at least the first plurality of BAM files along parallel processing paths. The plurality of segments may correspond to any given number of genomic subregions and may be selected based upon the number of processing cores used in the parallel processing.
Claims
exact text as granted — not AI-modifiedWhat is claimed:
1 - 64 . (canceled)
65 . A method, comprising:
receiving genetic sequence data associated with a subject; splitting the genetic sequence data into a set of segments, each segment from the set of segments including an overlap region that includes the same genetic sequence data as an overlap region of an adjacent segment from the set of segments; and processing the set of segments on a set of processors, including allocating each segment from the set of segments to a processing path from a set of parallel processing paths, each processing path from the set of parallel processing paths performing at least one of the following processing steps on one or more segments allocated to that processing path from the set of parallel processing paths: local realignment; deduplication; recalibration; or genotyping.
66 . The method of claim 65 , wherein the overlap region of each segment includes about 3 kilobases.
67 . The method of claim 65 , wherein a number of segments from the set of segments is based on a number of processors associated with the compute device.
68 . The method of claim 65 , wherein the processing includes the step of genotyping, the genotyping including identifying a presence of a variant, the genotyping further including, when the variant occurs in both an overlap region of a first segment from the set of segments and in an overlap region of a second segment from the set of segments that is adjacent to the first segment, assigning the variant to either the overlap region of the first segment or the overlap region of the second segment based on a location of the variant.
69 . The method of claim 65 , wherein the processing includes the step of genotyping, the genotyping including identifying a presence of a variant, the genotyping further including, when the variant occurs in both a first overlap region of a first segment and in a second overlap region of a second segment that is adjacent to the first segment, assigning the variant to either the first segment or the second segment based on a position of the variant relative to a boundary of the first segment and relative to a boundary of the second segment.
70 . The method of claim 65 , wherein the genetic sequence data specifies a plurality of chromosomes, and mitochondrial DNA.
71 . The method of claim 65 , wherein the processing further includes mapping reads to a reference genome.
72 . The method of claim 65 , wherein the processing further includes generating binary alignment map (BAM) files based on the set of segments.
73 . The method of claim 65 , wherein the set of parallel processing paths include at least 26 parallel processing paths.
74 . A non-transitory processor-readable medium storing code representing instructions to be executed by one or more processors from a set of processors, the code comprising code to cause the one or more processors from the set of processors to:
receive genetic sequence data associated with a subject; split the genetic sequence data into a set of segments, each segment from the set of segments including an overlap region that includes the same genetic sequence data as an overlap region of an adjacent segment from the set of segments; and process the set of segments by:
allocating each segment from the set of segments to a processing path from a set of parallel processing paths, each processing path from the set of parallel processing paths performing at least one processing step on one or more segments allocated to that processing path from the set of parallel processing paths; and
identifying the presence of a variant including, when the variant occurs in both an overlap region of a first segment from the set of segments and in an overlap region of a second segment from the set of segments that is adjacent to the first segment, assigning the variant to either the overlap region of the first segment or the overlap region of the second segment based on at least one of a location of the variant, or a position of the variant relative to a boundary of the first segment and relative to a boundary of the second segment.
75 . The non-transitory processor-readable medium of claim 74 , wherein the overlap region of each segment from the set of segments includes about 3 kilobases.
76 . The non-transitory processor-readable medium of claim 74 , wherein a number of segments in the set of segments is based on a number of processors in the set of processors.
77 . The non-transitory processor-readable medium of claim 74 , wherein the code to cause the one or more processors to process the set of segments includes code to cause each processing path from the set of parallel processing paths to perform at least one of the following processing steps on one or more segments allocated to that processing path from the set of parallel processing paths:
local realignment; deduplication; recalibration; or genotyping
78 . The non-transitory processor-readable medium of claim 74 , wherein the genetic sequence data specifies a plurality of chromosomes, and mitochondrial DNA.
79 . The non-transitory processor-readable medium of claim 74 , wherein the code to cause the one or more processors from the set of processors to process the set of segments includes code to cause the one or more processors to map reads to a reference genome, including dividing a mapping step among the set of parallel processing paths.
80 . The non-transitory processor-readable medium of claim 74 , wherein the code to cause the one or more processors from the set of processors to process the set of segments includes code to cause the one or more processors to generate binary alignment map (BAM) files based on the set of segments.
81 . The non-transitory processor-readable medium of claim 74 , wherein the set of parallel processing paths includes at least 26 parallel processing paths.Join the waitlist — get patent alerts
Track US2017220732A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.