Systems and Methods for Identifying Exon Junctions from Single Reads
Abstract
Systems and methods are used to identify an exon junction from a single read of a transcript. A transcript sample is interrogated and a read sequence is produced using a nucleic acid sequencer. A first exon sequence and a second exon sequence are obtained using the processor. The first exon sequence is mapped to a prefix of the read sequence using the processor. The second exon sequence is mapped to a suffix of the read sequence using the processor. A sum of a number of sequence elements of the first exon sequence that overlap the prefix of the read sequence, of a number of sequence elements of the second exon sequence that overlap the suffix of the read sequence, and of a constant is calculated using the processor. If the sum equals a length of the read sequence, a junction is identified in the read using the processor.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system for identifying a fusion junction in a transcript sample suspected of containing a gene fusion, the system comprising:
a nucleic acid sequencer configured to:
receive at least a portion of a fragment library including a plurality of nucleic acid fragments;
provide reagents for sequencing the nucleic acid fragments;
detect a plurality of signals during sequencing, the signals representative of a first sequence of at least one of the nucleic acid fragments;
a memory comprising a stored list of exon prefix sequences and a stored list of exon suffix sequences; a processor in communication with the nucleic acid sequencer and the memory, the processor configured to:
obtain a first read sequence based on the detected plurality of signals from the nucleic acid sequencer, the first read sequence corresponding to the first sequence,
map a first exon sequence, chosen from the stored list of exon suffix sequences, to a prefix of the first read sequence and map a second exon sequence, chosen from the stored list of exon prefix sequences, to a suffix of the first read sequence,
calculate a sum of a number of sequence elements of the first exon sequence that overlap the prefix of the first read sequence, a number of sequence elements of the second exon sequence that overlap the suffix of the first read sequence, and a constant, and
if the sum equals a length of the first read sequence, identify a fusion junction between exons associated with the first exon sequence and second exon sequence in the transcript sample, and identify the presence of a gene fusion in the transcript sample based on the identified fusion junction.
2 . The system of claim 1 , wherein the first exon sequence is a reverse sequence.
3 . The system of claim 1 , wherein the second exon sequence is a reverse sequence.
4 . The system of claim 1 , wherein the first exon sequence, the second exon sequence, and the first read sequence are monobase color-space sequences and the constant is 0.
5 . The system of claim 1 , wherein the first exon sequence, the second exon sequence, and the first read sequence are dibase color-space sequences and the constant is 1.
6 . The system of claim 1 , wherein the processor maps the first exon sequence to a prefix of the first read sequence by at least a minimum number of sequence elements.
7 . The system of claim 6 , wherein the minimum number of sequence elements is defined by a user.
8 . The system of claim 1 , wherein the exon prefix sequences and exon suffix sequences of the stored lists comprise sequences of a length ranging from 10 to 40 bases.
9 . A method for identifying a fusion junction in a transcript sample suspected of containing a gene fusion, the method comprising:
preparing a fragment library from nucleic acids isolated from the transcript sample; providing at least a portion of the fragment library to a sequencing instrument; detecting a plurality of signals, at least some of which are representative of a sequence of one of the nucleic acid fragments of the fragment library; using a processor to:
generate a first read sequence representative of the nucleic acid fragment from the plurality of signals;
retrieve a first exon sequence from a list of exon suffix sequences stored in a computer memory;
retrieve a second exon sequence from a list of exon prefix sequences stored in a computer memory;
map the first exon sequence to a prefix of the first read sequence;
map the second exon sequence to a suffix of the first read sequence;
calculate a sum of a number of sequence elements of the first exon sequence that overlap the prefix of the first read sequence, a number of sequence elements of the second exon sequence that overlap the suffix of the first read sequence, and a constant;
if the sum equals a length of the first read sequence, identify a fusion junction between exons associated with the first exon sequence and second exon sequence in the transcript sample using the processor; and
identify the presence of a gene fusion in the transcript sample based on the identified fusion junction.
10 . The method of claim 9 , further comprising:
generating a second read sequence from the plurality of signals; mapping the first exon sequence to a prefix of the second read sequence, and mapping the second exon sequence to a suffix of the second read sequence.
11 . The method of claim 10 , wherein the second read sequence is a paired end read sequence.
12 . The method of claim 9 , further comprising calculating a confidence value for the junction.
13 . The method of claim 11 , wherein the confidence value depends on a number of unique read sequences corresponding to the junction.
14 . The method of claim 9 , wherein the exon prefix sequences and exon suffix sequences of the stored lists comprise sequences of a length ranging from 10 to 40 bases.
15 . A computer program product, comprising a non-transitory computer-readable storage medium whose contents include a program with instructions being executed on a processor so as to perform a method for identifying a fusion junction in a transcript sample suspected of containing a gene fusion, the instructions comprising:
instructions to receive at least a portion of a fragment library including a plurality of nucleic acid fragments into a sequencing instrument; instructions to provide reagents for sequencing the nucleic acid fragments; instructions to detect a plurality of signals during sequencing, at least a portion of the signals representative of a sequence of at least one of the nucleic acid fragments; instructions to determine a first read sequence representative of the at least one nucleic acid fragment using the plurality of signals; instructions to store a list of exon prefix sequences and a list of exon suffix sequences; instructions to obtain a first exon sequence, chosen from the stored list of exon suffix sequences, and a second exon sequence, chosen from the stored list of exon prefix sequences; instructions to map the first exon sequence to a prefix of the first read sequence; instructions to map the second exon sequence to a suffix of the first read sequence; instructions to calculate a sum of a number of sequence elements of the first exon sequence that overlap the prefix of the first read sequence, a number of sequence elements of the second exon sequence that overlap the suffix of the first read sequence, and a constant; instructions to identify a fusion junction between exons associated with the first exon sequence and second exon sequence when the sum equals a length of the first read sequence; and instructions to identify the presence of a gene fusion in the transcript sample based on the identified fusion junction.
16 . The computer program product of claim 15 , wherein the first exon sequence, the second exon sequence, and the first read sequence are monobase color-space sequences and the constant is 0.
17 . The computer program product of claim 15 , wherein the first exon sequence, the second exon sequence, and the first read sequence are dibase color-space sequences and the constant is 1.
18 . The computer program product of claim 15 , wherein the instructions further comprise:
instructions to generate a second read sequence from the plurality of signals; instructions to map the first exon sequence to a prefix of the second read sequence, and instructions to map the second exon sequence to a suffix of the second read sequence.
19 . The computer program product of claim 18 , wherein the second read sequence is a paired end read sequence.
20 . The computer program product of claim 15 , wherein the instructions further comprise instructions to calculate a confidence value for the junction.Join the waitlist — get patent alerts
Track US2018276338A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.