US2012035857A1PendingUtilityA1

Computer-implemented biological sequence identifier system and method

Individually held — no corporate assignee on recordPriority: Jul 2, 2004Filed: Aug 17, 2011Published: Feb 9, 2012
Est. expiryJul 2, 2024(expired)· nominal 20-yr term from priority
G16B 30/10G16B 30/00C12Q 1/6874C12Q 1/689C12Q 1/6888C12Q 1/701C12Q 1/6893
57
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A computer-implemented biological sequence identifier (CIBSI) system and method for selecting a subsequence from biological sequence data according to at least one selection parameter. The at least one selection parameter corresponds to a likelihood of returning a meaningful result from a similarity search.

Claims

exact text as granted — not AI-modified
1 . A computer-implemented method for selecting a biological subsequence for input to a query for identification of a predetermined biological sequence, comprising steps of:
 selecting with a processor-implemented process a subsequence from biological sequence data stored in memory; and   submitting the subsequence in a query to identify the predetermined biological sequence with a first predetermined confidence level;
 wherein the first predetermined confidence level is above a selected threshold. 
   
     
     
         2 . The computer implemented method of  claim 1 , further comprising:
 storing the biological sequence data in one of a PASTA, MSF, GCG, Clustal, BLC, PIR, MSP, PFAM, POSTAL and JNET format.   
     
     
         3 . The computer-implemented method of  claim 2 , further comprising:
 determining whether the biological sequence data corresponds to one of a biological sequence or a control sequence.   
     
     
         4 . The computer-implemented method of  claim 1 , wherein the selecting step comprises:
 selecting a window size parameter corresponding to a number of base calls in the biological sequence data; and   calculating a percentage of valid base calls contained within a viewing window of the biological sequence data, the size of the window corresponding to the window size parameter selected in the selecting step.   
     
     
         5 . The computer-implemented method of  claim 4 , wherein the selecting step comprises:
 sliding the viewing window to another number of base calls in the biological sequence when the percentage calculated in the calculating step does not satisfy a predetermined threshold; and   calculating a percentage of valid base calls contained within the another number of base calls in the biological sequence.   
     
     
         6 . The computer-implemented method of  claim 4 , wherein the selecting step comprises:
 selecting the subsequence of base calls within a viewing window as the subsequence submitted in the query when the calculated percentage satisfies a predetermined threshold.   
     
     
         7 . The computer-implemented method of  claim 6 , further comprising:
 trimming invalid base calls from the selected subsequence of base calls before the selected subsequence is submitted in the submitting step.   
     
     
         8 . The computer-implemented method of  claim 1 , further comprising:
 comparing the subsequence with a plurality of predetermined sequences; and   generating comparison results corresponding to at least one of said predetermined sequences.   
     
     
         9 . The computer-implemented method of  claim 8 , wherein comparison results from the comparing step include a statistical value indicating a predetermined level of correspondence between the subsequence and the at least one of said predetermined sequences. 
     
     
         10 . The computer-implemented method of  claim 1 , further comprising:
 incorporating intensity data with the biological sequence data; and   estimating a concentration of at least one target sequence.   
     
     
         11 . The computer-implemented method of  claim 10 , further comprising:
 detecting at least two subsequences from the biological sequence data according to at least one selection parameter; and   detecting at least one of a mixture and a recombination event.   
     
     
         12 . The computer-implemented method of  claim 11 , wherein the at least two subsequences correspond to different regions of a microarray. 
     
     
         13 . The computer-implemented method of  claim 10 , further comprising:
 distinguishing between a mixture of similar sequences and a recombination between different sequences;   wherein the similar sequences have a predetermined level of similarity.   
     
     
         14 . The computer-implemented method of  claim 10 , further comprising:
 distinguishing between a mixture and a recombination event, including evaluating a first signal from a first region of the microarray and a second signal from a second region of the microarray, and   comparing the first signal to the second signal to generate at least one distinction parameter, the at least one distinction parameter corresponding to a probability the first signal and the second signal indicate one of a mixture and a recombination event.   
     
     
         15 . The computer-implemented method of  claim 1 , further comprising:
 identifying at least one consensus sequence corresponding to a. plurality of test sequences;   selecting the subsequence from the at least one consensus sequence;   comparing the at least one subsequence with at least one predetermined sequence;   generating a comparison result;   calculating a difference between the comparison result and the plurality of test sequences; and   generating at least one candidate consensus sequence.   
     
     
         16 . The computer-implemented method of  claim 15 , further comprising:
 producing a microarray probe according to the at least one candidate consensus sequence.   
     
     
         17 . The computer implemented method of  claim 15 , further comprising:
 modifying the at least one consensus sequence according to a patch parameter, the patch parameter corresponding to at least a portion of at least one of the plurality of test sequences.   
     
     
         18 . The computer-implemented method of  claim 15 , further comprising:
 simulating a hybridization between the at least one candidate consensus sequence and the plurality of test sequences according to at least one hybridization parameter.   
     
     
         19 . The computer-implemented method of  claim 15 , wherein the biological sequence data includes at least one of a nucleic acid, a transcriptional monomer, a transcription product, DNA, and RNA. 
     
     
         20 . The computer-implemented method of  claim 1 , wherein the biological sequence data includes at least one of a gap and an ambiguous subsequence. 
     
     
         21 . The computer-implemented method of  claim 1 , further comprising:
 calculating a relative position of the biological sequence data, wherein the biological sequence data includes at least one of an amino acid and a protein.   
     
     
         22 . The computer-implemented method of  claim 1 , further comprising:
 obtaining the biological sequence data by at least one of manual Sanger sequencing, automated Sanger sequencing, shotgun sequencing, conventional microarrays, resequencing microarrays, microelectrophoretic sequencing, sequencing by hybridization (SBH), Edman degradation, Cyclic-array sequencing on amplified molecules, Cyclic-array sequencing on single molecules and nanopore sequencing.   
     
     
         23 . The computer-implemented method of  claim 1 , wherein the biological sequence data is at least one of a nucleotide sequence and a protein sequence. 
     
     
         24 . A computer readable storage medium configured to store computer readable instructions for execution on a computer, the computer readable instructions, when executed by the computer, configured to perform the method of identifying a predetermined biological sequence comprising the steps of:
 selecting with a processor implemented process a. subsequence from biological sequence data stored in a memory; and   submitting the subsequence in a query to identify the predetermined biological sequence with a first predetermined confidence level;
 wherein the first confidence level is above a selection threshold. 
   
     
     
         25 . An apparatus for selecting a biological subsequence for input to a query for identification of a. predetermined biological sequence, comprising:
 means for selecting a subsequence from biological sequence data stored in a memory; and   means for submitting the subsequence in a query to identify the predetermined biological sequence with a first predetermined confidence level;
 wherein the first confidence level is above a selection threshold.

Join the waitlist — get patent alerts

Track US2012035857A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.