US2016103956A1PendingUtilityA1

Biological sequence variant characterization

Assignee: IBMPriority: Oct 10, 2014Filed: Jul 7, 2015Published: Apr 14, 2016
Est. expiryOct 10, 2034(~8.2 yrs left)· nominal 20-yr term from priority
G16B 20/00G06F 19/22G16B 30/00G16B 20/20
41
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Short fixed length sub-sequences, defined as reference sub-sequences, are extracted from a collection of reference sequences, and an index is constructed showing which short fixed length reference sub-sequence occurs in which reference sequences. Short fixed length sub-sequences, the same length as the reference sub-sequences and defined as source sub-sequences, are extracted from a collection of source sequences derived from a sample for which the signature is to be determined, and the short fixed length source sub-sequences are compiled to determine the frequency of each within the collection. The presence or absence of source sub-sequences in combination with the index is used to infer the presence or absence of reference sequences from the reference collection.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for determining a signature from biological sequence data, comprising:
 extracting short fixed length sub-sequences, defined as reference sub-sequences, from a collection of reference sequences, and constructing an index showing which short fixed length reference sub-sequence occurs in which reference sequences;   extracting short fixed length sub-sequences, the same length as the reference sub-sequences and defined as source sub-sequences, from a collection of source sequences derived from a sample for which the signature is to be determined, and compiling the short fixed length source sub-sequences to determine the frequency of each within the collection; and   using the presence or absence of source sub-sequences in combination with the index to infer the presence or absence of reference sequences from the reference collection;   wherein one or more of the above steps are performed in accordance with a processor and a memory.   
     
     
         2 . The method of  claim 1 , wherein when a member of a subset of the reference collection is expected but not found to be present, the initial and final short fixed length reference sub-sequences of the members of the subset of the reference sequences are used, along with overlaps between the short fixed length source sub-sequences, to derive a novel member of that subset of the reference sequences. 
     
     
         3 . The method of  claim 1 , further comprising using at least one of the frequency of the extracted short fixed length source sub-sequences and overlap relationships between the short fixed length source sub-sequences to eliminate short fixed length source sub-sequences that are due to measurement error. 
     
     
         4 . The method of  claim 1 , wherein the biological sequence data is classified as at least one of DNA, RNA, and proteins. 
     
     
         5 . The method of  claim 1 , wherein the short fixed length sub-sequences comprise k-mers. 
     
     
         6 . The method of  claim 1 , wherein the sample from which the source sub-sequences are derived comprises one of next-generation sequences, high-throughput sequences, whole-genome sequences, whole-exome sequences, transcriptomes, and metagenomes. 
     
     
         7 . The method of  claim 1 , wherein the reference sequence collection comprises an allelic genotyping scheme. 
     
     
         8 . The method of  claim 7 , wherein the allelic genotyping scheme comprises one of a multi-locus sequence typing and whole-genome sequence typing. 
     
     
         9 . The method of  claim 1 , wherein sub-sequence overlaps from which novel sequence variants are identified are in the form of a de Bruijn graph. 
     
     
         10 . A method comprising:
 collating short fixed length sequences from each position in a reference subsequence to produce an index;   collating short fixed length sequences from experimentally derived sequences, producing a list of short fixed length sequences with their associated frequency;   discarding short fixed length sequences with low frequency; and   for each of the remaining short fixed length sequences in the index, making a record of which ones are present, such that when the indexed short fixed length sequences are present, a conclusion is made that the reference subsequence is present in the experimentally derived sequences;   wherein one or more of the above steps are performed in accordance with a processor and a memory.   
     
     
         11 . The method of  claim 10 , wherein the index comprises a matrix with a row for each short fixed length sequence in the collection of reference subsequences, and a column for each reference sequence, and for each column, a first value is stored in the rows corresponding to the short fixed length sequences that are present in the reference subsequence for that column. 
     
     
         12 . The method of  claim 11 , wherein the short fixed length sequences are extracted from the experimentally derived sub-sequences. 
     
     
         13 . The method of  claim 12 , wherein for each short fixed length sequence, the corresponding row of the matrix is consulted and for each reference subsequence corresponding to a column containing the first value, the presence of the short fixed length sequence is recorded and, after processing the short fixed length sequences, any reference subsequence for which its short fixed length sequences were observed is reported as being present in the experimentally derived data.

Join the waitlist — get patent alerts

Track US2016103956A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.