US2018276338A1PendingUtilityA1

Systems and Methods for Identifying Exon Junctions from Single Reads

Assignee: LIFE TECHNOLOGIES CORPPriority: Apr 30, 2010Filed: Mar 22, 2018Published: Sep 27, 2018
Est. expiryApr 30, 2030(~3.8 yrs left)· nominal 20-yr term from priority
G06F 19/22G16B 30/10G16B 30/00
55
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods are used to identify an exon junction from a single read of a transcript. A transcript sample is interrogated and a read sequence is produced using a nucleic acid sequencer. A first exon sequence and a second exon sequence are obtained using the processor. The first exon sequence is mapped to a prefix of the read sequence using the processor. The second exon sequence is mapped to a suffix of the read sequence using the processor. A sum of a number of sequence elements of the first exon sequence that overlap the prefix of the read sequence, of a number of sequence elements of the second exon sequence that overlap the suffix of the read sequence, and of a constant is calculated using the processor. If the sum equals a length of the read sequence, a junction is identified in the read using the processor.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system for identifying a fusion junction in a transcript sample suspected of containing a gene fusion, the system comprising:
 a nucleic acid sequencer configured to:
 receive at least a portion of a fragment library including a plurality of nucleic acid fragments; 
 provide reagents for sequencing the nucleic acid fragments; 
 detect a plurality of signals during sequencing, the signals representative of a first sequence of at least one of the nucleic acid fragments; 
   a memory comprising a stored list of exon prefix sequences and a stored list of exon suffix sequences;   a processor in communication with the nucleic acid sequencer and the memory, the processor configured to:
 obtain a first read sequence based on the detected plurality of signals from the nucleic acid sequencer, the first read sequence corresponding to the first sequence, 
 map a first exon sequence, chosen from the stored list of exon suffix sequences, to a prefix of the first read sequence and map a second exon sequence, chosen from the stored list of exon prefix sequences, to a suffix of the first read sequence, 
 calculate a sum of a number of sequence elements of the first exon sequence that overlap the prefix of the first read sequence, a number of sequence elements of the second exon sequence that overlap the suffix of the first read sequence, and a constant, and 
 if the sum equals a length of the first read sequence, identify a fusion junction between exons associated with the first exon sequence and second exon sequence in the transcript sample, and identify the presence of a gene fusion in the transcript sample based on the identified fusion junction. 
   
     
     
         2 . The system of  claim 1 , wherein the first exon sequence is a reverse sequence. 
     
     
         3 . The system of  claim 1 , wherein the second exon sequence is a reverse sequence. 
     
     
         4 . The system of  claim 1 , wherein the first exon sequence, the second exon sequence, and the first read sequence are monobase color-space sequences and the constant is 0. 
     
     
         5 . The system of  claim 1 , wherein the first exon sequence, the second exon sequence, and the first read sequence are dibase color-space sequences and the constant is 1. 
     
     
         6 . The system of  claim 1 , wherein the processor maps the first exon sequence to a prefix of the first read sequence by at least a minimum number of sequence elements. 
     
     
         7 . The system of  claim 6 , wherein the minimum number of sequence elements is defined by a user. 
     
     
         8 . The system of  claim 1 , wherein the exon prefix sequences and exon suffix sequences of the stored lists comprise sequences of a length ranging from 10 to 40 bases. 
     
     
         9 . A method for identifying a fusion junction in a transcript sample suspected of containing a gene fusion, the method comprising:
 preparing a fragment library from nucleic acids isolated from the transcript sample;   providing at least a portion of the fragment library to a sequencing instrument;   detecting a plurality of signals, at least some of which are representative of a sequence of one of the nucleic acid fragments of the fragment library;   using a processor to:
 generate a first read sequence representative of the nucleic acid fragment from the plurality of signals; 
 retrieve a first exon sequence from a list of exon suffix sequences stored in a computer memory; 
 retrieve a second exon sequence from a list of exon prefix sequences stored in a computer memory; 
 map the first exon sequence to a prefix of the first read sequence; 
 map the second exon sequence to a suffix of the first read sequence; 
 calculate a sum of a number of sequence elements of the first exon sequence that overlap the prefix of the first read sequence, a number of sequence elements of the second exon sequence that overlap the suffix of the first read sequence, and a constant; 
 if the sum equals a length of the first read sequence, identify a fusion junction between exons associated with the first exon sequence and second exon sequence in the transcript sample using the processor; and 
 identify the presence of a gene fusion in the transcript sample based on the identified fusion junction. 
   
     
     
         10 . The method of  claim 9 , further comprising:
 generating a second read sequence from the plurality of signals;   mapping the first exon sequence to a prefix of the second read sequence, and   mapping the second exon sequence to a suffix of the second read sequence.   
     
     
         11 . The method of  claim 10 , wherein the second read sequence is a paired end read sequence. 
     
     
         12 . The method of  claim 9 , further comprising calculating a confidence value for the junction. 
     
     
         13 . The method of  claim 11 , wherein the confidence value depends on a number of unique read sequences corresponding to the junction. 
     
     
         14 . The method of  claim 9 , wherein the exon prefix sequences and exon suffix sequences of the stored lists comprise sequences of a length ranging from 10 to 40 bases. 
     
     
         15 . A computer program product, comprising a non-transitory computer-readable storage medium whose contents include a program with instructions being executed on a processor so as to perform a method for identifying a fusion junction in a transcript sample suspected of containing a gene fusion, the instructions comprising:
 instructions to receive at least a portion of a fragment library including a plurality of nucleic acid fragments into a sequencing instrument;   instructions to provide reagents for sequencing the nucleic acid fragments;   instructions to detect a plurality of signals during sequencing, at least a portion of the signals representative of a sequence of at least one of the nucleic acid fragments;   instructions to determine a first read sequence representative of the at least one nucleic acid fragment using the plurality of signals;   instructions to store a list of exon prefix sequences and a list of exon suffix sequences;   instructions to obtain a first exon sequence, chosen from the stored list of exon suffix sequences, and a second exon sequence, chosen from the stored list of exon prefix sequences;   instructions to map the first exon sequence to a prefix of the first read sequence;   instructions to map the second exon sequence to a suffix of the first read sequence;   instructions to calculate a sum of a number of sequence elements of the first exon sequence that overlap the prefix of the first read sequence, a number of sequence elements of the second exon sequence that overlap the suffix of the first read sequence, and a constant;   instructions to identify a fusion junction between exons associated with the first exon sequence and second exon sequence when the sum equals a length of the first read sequence; and   instructions to identify the presence of a gene fusion in the transcript sample based on the identified fusion junction.   
     
     
         16 . The computer program product of  claim 15 , wherein the first exon sequence, the second exon sequence, and the first read sequence are monobase color-space sequences and the constant is 0. 
     
     
         17 . The computer program product of  claim 15 , wherein the first exon sequence, the second exon sequence, and the first read sequence are dibase color-space sequences and the constant is 1. 
     
     
         18 . The computer program product of  claim 15 , wherein the instructions further comprise:
 instructions to generate a second read sequence from the plurality of signals;   instructions to map the first exon sequence to a prefix of the second read sequence, and   instructions to map the second exon sequence to a suffix of the second read sequence.   
     
     
         19 . The computer program product of  claim 18 , wherein the second read sequence is a paired end read sequence. 
     
     
         20 . The computer program product of  claim 15 , wherein the instructions further comprise instructions to calculate a confidence value for the junction.

Join the waitlist — get patent alerts

Track US2018276338A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.