Parallel-processing systems and methods for highly scalable analysis of biological sequence data
Abstract
An apparatus includes a memory configured to store a sequence that includes an estimation of a biological sequence. The sequence includes a set of elements. The apparatus also includes an assignment module implemented in a hardware processor. The assignment module is configured to receive the sequence from the memory, and assign each element to at least one segment from a set of segments, including, when an element maps to at least a first segment and a second segment, assigning the element set of segments specific to that hardware processor, and substantially simultaneous with the remaining hardware processors, remove at least a portion of duplicate elements in that segment to generate a deduplicated segment. Reorder the elements in the deduplicated segment to generate a realigned segment that has a reduced likelihood for alignment errors
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An apparatus, comprising:
a memory configured to store a sequence, the sequence including an estimation of a biological sequence, the sequence including a plurality of elements; and a plurality of hardware processors operatively coupled to the memory, each hardware processor from the plurality of hardware processors configured to implement a segment processing module, an assignment module implemented in a hardware processor from the plurality of hardware processors, the assignment module configured to:
receive the sequence from the memory, and
assign each element from the plurality of elements to at least one segment from a plurality of segments, including, when an element from the plurality of elements maps to at least a first segment and a second segment from the plurality of segments, assigning the element from the plurality of elements to both the first segment and the second segment,
the segment processing module for each hardware processor from the plurality of hardware processors operatively coupled to the assignment module, the segment processing module for each hardware processor from the plurality of hardware processors configured to, for each segment from a set of segments specific to that hardware processor and from the plurality of segments, and substantially simultaneous with the remaining hardware processors from the plurality of hardware processors:
remove at least a portion of duplicate elements in that segment from that set of segments to generate a deduplicated segment; and
reorder the elements in the deduplicated segment to generate a realigned segment, the realigned segment having a reduced likelihood for alignment errors than the deduplicated segment.
2 . The apparatus of claim 1 , wherein the sequence is a target sequence, the assignment module further configured to receive the target sequence by:
receiving a first sequence, the first sequence including a forward estimation of the biological sequence; receiving a second sequence, the second sequence including a reverse estimation of the biological sequence; generating a paired sequence based on the first sequence and the second sequence; and aligning the paired sequence with a reference sequence to generate the target sequence.
3 . The apparatus of claim 1 , wherein the sequence is in a binary alignment/map (BAM) format or in a FASTQ format.
4 . The apparatus of claim 1 , wherein the biological sequence is one of a deoxyribonucleic acid (DNA) sequence or a ribonucleic (RNA) sequence.
5 . A method, comprising:
receiving a sequence, the sequence including an estimation of a biological sequence, the sequence including a plurality of elements; assigning each element from the plurality of elements to at least one segment from a plurality of segments, including, when an element from the plurality of elements maps to both a first segment and a second segment from the plurality of segments, assigning the element from the plurality of elements to both the first segment and the second segment; and for each segment from the plurality of segments:
removing at least a portion of duplicate elements in the segment to generate a deduplicated segment;
reordering the elements in the deduplicated segment to generate a realigned segment, the realigned segment having a reduced likelihood for alignment errors than the deduplicated segment; and
transmitting the realigned segment to one or more of a storage module and a genotyping module.
6 . The method of claim 5 , wherein the sequence is a target sequence, the receiving the target sequence including:
receiving a first sequence, the first sequence including a forward estimation of the biological sequence; receiving a second sequence, the second sequence including a reverse estimation of the biological sequence; generating a paired sequence based on the first sequence and the second sequence; and aligning the paired sequence with a reference sequence to generate the target sequence.
7 . The method of claim 5 , further comprising:
prior to the assigning, splitting the sequence into a plurality of subsequences, the assigning including assigning each subsequence from the plurality of subsequences to at least one segment from the plurality of segments; and for each segment from the plurality of segments, subsequent to the assigning and prior to the removing, combining subsequences within the segment.
8 . The method of claim 5 , further comprising, when an element from the plurality of elements maps to the first segment and the second segment from the plurality of segments, assigning the first segment and the second segment to an intersegmental sequence.
9 . The method of claim 5 , wherein the sequence is in a binary alignment/map (BAM) format or in a FASTQ format.
10 . The method of claim 5 , wherein the sequence includes quality score information.
11 . The method of claim 5 , wherein each element from the plurality of elements includes a read pair.
12 . The method of claim 5 , wherein the biological sequence is one of a deoxyribonucleic acid (DNA) sequence or a ribonucleic (RNA) sequence.
13 . The method of claim 5 , wherein each segment from the plurality of segments includes a portion overlapping a portion of at least one remaining segment from the plurality of segments.
14 . The method of claim 5 , wherein each segment from the plurality of segments includes a portion having a first size overlapping a portion of at least one remaining segment from the plurality of segments,
the deduplicated segment associated with each segment from the plurality of segments including a portion having a second size overlapping a portion of the deduplicated segment associated with a remaining segment from the plurality of segments, the second size being smaller than the first size.
15 . The method of claim 5 , wherein each segment from the plurality of segments includes a portion having a first size overlapping a portion of at least one remaining segment from the plurality of segments,
the deduplicated segment associated with each segment from the plurality of segments including a portion having a second size overlapping a portion of the deduplicated segment associated with a remaining segment from the plurality of segments, the realigned segment associated with each segment from the plurality of segments including a portion having a third size overlapping a portion of the realigned segment associated with a remaining segment from the plurality of segments, the second size being smaller than the first size, the third size being smaller than the second size.
16 . An apparatus, comprising:
an assignment module, implemented in a memory or a processor, configured to:
receive a sequence, the sequence including an estimation of a biological sequence, the sequence including a plurality of elements;
assign each element from the plurality of elements to at least one segment from a plurality of segments, including, when an element from the plurality of elements maps to at least a first segment and a second segment from the plurality of segments, assigning the element from the plurality of elements to both the first segment and the second segment; and
a segment processing module operatively coupled to the assignment module, the segment processing module configured to, for each segment from the plurality of segments:
remove at least a portion of duplicate elements in the segment generate a deduplicated segment; and
reorder the elements in the deduplicated segment to generate a realigned segment, the realigned segment having a reduced likelihood for alignment errors than the deduplicated segment,
the segment processing module further configured to execute the removing and the reordering for at least two segments from the plurality of segments in a substantially simultaneous manner.
17 . The apparatus of claim 16 , wherein the segment processing module is further configured to execute the removing and the reordering for the plurality of segments in a substantially simultaneous manner.
18 . The apparatus of claim 16 , wherein the sequence is a target sequence, the assignment module further configured to receive the target sequence by:
receiving a first sequence, the first sequence including a forward estimation of the biological sequence; receiving a second sequence, the second sequence including a reverse estimation of the biological sequence; generating a paired sequence based on the first sequence and the second sequence; and aligning the paired sequence with a reference sequence to generate the target sequence.
19 . The apparatus of claim 16 , further comprising a parallelization module configured to, prior to the assigning by the assignment module, split the sequence into a plurality of subsequences,
the assignment module configured to assign by assigning each subsequence to at least one segment from the plurality of segments, the assignment module further configured to execute the assigning for at least two segments from the plurality of segments in a substantially simultaneous manner, the parallelization module further configured to, for each segment from the plurality of segments, subsequent to the assigning by the assignment module and prior to the removing by the segment processing module, combine subsequences within that segment.
20 . The apparatus of claim 16 , wherein each segment from the plurality of segments includes a portion having a first size overlapping a portion of at least one remaining segment from the plurality of segments,
the deduplicated segment associated with each segment from the plurality of segments including a portion having a second size overlapping a portion of the deduplicated segment associated with a remaining segment from the plurality of segments, the realigned segment associated with each segment from the plurality of segments including a portion having a third size overlapping a portion of the realigned segment associated with a remaining segment from the plurality of segments, the second size being smaller than the first size, the third size being smaller than the second size.
21 . The apparatus of claim 16 , wherein the sequence is in a binary alignment/map (BAM) format or in a FASTQ format.
22 . The apparatus of claim 16 , wherein the biological sequence is one of a deoxyribonucleic acid (DNA) sequence or a ribonucleic (RNA) sequence.Join the waitlist — get patent alerts
Track US2017316154A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.