US2020303039A1PendingUtilityA1

Methods and systems for sequence calling

Assignee: ULTIMA GENOMICS INCPriority: Oct 26, 2017Filed: Apr 10, 2020Published: Sep 24, 2020
Est. expiryOct 26, 2037(~11.2 yrs left)· nominal 20-yr term from priority
C12Q 1/6869G16B 30/10G16B 45/00G16B 40/10C12Q 1/6806
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present disclosure provides methods and systems for accurate and efficient context-aware base calling of sequences. In an aspect, disclosed herein is a method for sequencing a nucleic acid molecule, comprising: (a) sequencing the nucleic acid molecule to generate a plurality of sequence signals; and (b) determining base calls of the nucleic acid molecule based at least in part on (i) the plurality of sequence signals and (ii) quantified context dependency for at least a portion of the plurality of sequence signals.

Claims

exact text as granted — not AI-modified
1 - 97 . (canceled) 
     
     
         98 . A method for processing a plurality of sequence signals, comprising:
 (a) sequencing a nucleic acid sample to provide said plurality of sequence signals and a plurality of imputed sequences, wherein said plurality of imputed sequences comprises a plurality of homopolymer sequences;   (b) truncating said plurality of homopolymer sequences to provide a plurality of HpN truncated sequences comprising truncated homopolymer sequences of N bases in length;   (c) aligning said plurality of HpN truncated sequences to a truncated reference sequence, wherein said truncated reference sequence comprises one or more homopolymer sequences of N bases in length; and   (d) generating a consensus sequence based at least in part on said plurality of HpN truncated sequences aligned to said truncated reference sequence, wherein said consensus sequence comprises one or more homopolymer sequences of N bases in length.   
     
     
         99 . The method of  claim 98 , further comprising truncating a reference sequence comprising one or more homopolymer sequences each comprising at least N bases to provide said truncated reference sequence. 
     
     
         100 . The method of  claim 98 , wherein generating said consensus sequence in (d) is based at least in part on at least a subset of said plurality of sequence signals. 
     
     
         101 . The method of  claim 98 , further comprising determining a length estimation error of said one or more homopolymer sequences. 
     
     
         102 . The method of  claim 101 , wherein said length estimation error comprises a confidence interval for a length of said one or more homopolymer sequences. 
     
     
         103 . The method of  claim 101 , wherein determining said length estimation error is based at least in part on said plurality of HpN truncated sequences aligned to said truncated reference sequence. 
     
     
         104 . The method of  claim 98 , further comprising determining lengths of at least a subset of said plurality of homopolymer sequences based at least in part on clustering of said plurality of homopolymer sequences or said plurality of sequence signals associated with said plurality of homopolymer sequences. 
     
     
         105 . The method of  claim 98 , wherein N is 2 bases. 
     
     
         106 . The method of  claim 105 , wherein N is 3 bases. 
     
     
         107 . The method of  claim 98 , wherein said nucleic acid sample is prepared using fluorescently labeled nucleotides. 
     
     
         108 . A method for quantifying a context dependency of a plurality of sequence signals, the method comprising:
 (a) sequencing a nucleic acid sample to provide said plurality of sequence signals and a plurality of imputed sequences, wherein said nucleic acid sample comprises a known sequence, and wherein said plurality of imputed sequences comprises a plurality of homopolymer sequences;   (b) truncating said plurality of homopolymer sequences to provide a plurality of HpN truncated sequences comprising truncated homopolymer sequences of N bases in length;   (c) aligning said plurality of HpN truncated sequences to a truncated reference sequence, wherein said truncated reference sequence comprises one or more homopolymer sequences of N bases in length; and   (d) quantifying said context dependency of said plurality of sequence signals based at least in part on (i) said known sequence and (ii) said plurality of HpN truncated sequences aligned to said truncated reference sequence.   
     
     
         109 . The method of  claim 108 , further comprising:
 (e) sequencing an additional nucleic acid sample comprising an unknown nucleic acid sequence to provide an additional plurality of sequence signals and an additional plurality of imputed sequences, wherein said additional plurality of imputed sequences comprises an additional plurality of homopolymer sequences;   (f) truncating said additional plurality of homopolymer sequences to provide an additional plurality of HpN truncated sequences comprising additional truncated homopolymer sequences of N bases in length;   (g) aligning said additional plurality of HpN truncated sequences to said truncated reference sequence; and   (h) determining lengths of a least a subset of said additional plurality of homopolymer sequences based at least in part on (i) said context dependency and (ii) said plurality of HpN truncated sequences aligned to said truncated reference sequence.   
     
     
         110 . The method of  claim 108 , wherein said context dependency is associated with a given context. 
     
     
         111 . The method of  claim 110 , wherein said given context is an n-base context, wherein ‘n’ is a number greater than or equal to 5. 
     
     
         112 . The method of  claim 108 , wherein quantifying said context dependency comprises establishing a context specific mapping between signal amplitudes and homopolymer length for each of a plurality of loci. 
     
     
         113 . The method of  claim 108 , further comprising truncating a reference sequence comprising one or more homopolymer sequences each comprising at least N bases to provide said truncated reference sequence. 
     
     
         114 . The method of  claim 108 , further comprising determining lengths of at least a subset of said plurality of homopolymer sequences based at least in part on clustering of said plurality of homopolymer sequences or said plurality of sequence signals associated with said plurality of homopolymer sequences. 
     
     
         115 . The method of  claim 108 , wherein N is 2 bases. 
     
     
         116 . The method of  claim 115 , wherein N is 3 bases. 
     
     
         117 . The method of  claim 108 , wherein said nucleic acid sample is prepared using fluorescently labeled nucleotides. 
     
     
         118 . A system for quantifying a context dependency of a plurality of sequence signals, comprising:
 a database for storing said plurality of sequence signals and a plurality of imputed sequences, wherein said plurality of imputed sequences comprises a plurality of homopolymer sequences, and wherein said plurality of sequence signals and said plurality of imputed sequences correspond to a nucleic acid sample comprising a known sequence; and   one or more computer processors coupled to said database, wherein said one or more computer processors are individually or collectively programmed to:   (a) truncate said plurality of homopolymer sequences to provide a plurality of HpN truncated sequences comprising truncated homopolymer sequences of N bases in length;   (b) align said plurality of HpN truncated sequences to a truncated reference sequence, wherein said truncated reference sequence comprises one or more homopolymer sequences of N bases in length; and   (c) quantify said context dependency of said plurality of sequence signals based at least in part on (i) said known sequence and (ii) said plurality of HpN truncated sequences aligned to said truncated reference sequence.   
     
     
         119 . The system of  claim 118 , wherein said (i) database is further for storing an additional plurality of sequence signals and an additional plurality of imputed sequences, wherein said additional plurality of imputed sequences comprises an additional plurality of homopolymer sequences, and wherein said additional plurality of sequence signals and said additional plurality of imputed sequences correspond to an additional nucleic acid sample comprising an unknown sequence; and (ii) said one or more computer processors are further individually or collectively programmed to:
 (a) truncate said additional plurality of homopolymer sequences to provide an additional plurality of HpN truncated sequences comprising additional truncated homopolymer sequences of N bases in length; 
 (b) align said additional plurality of HpN truncated sequences to said truncated reference sequence; and 
 (c) determine lengths of a least a subset of said additional plurality of homopolymer sequences based at least in part on (i) said context dependency and (ii) said plurality of HpN truncated sequences aligned to said truncated reference sequence. 
 
     
     
         120 . The system of  claim 118 , wherein said one or more computer processors are further individually or collectively programmed to truncate a reference sequence comprising one or more homopolymer sequences each comprising at least N bases to provide said truncated reference sequence. 
     
     
         121 . The system of  claim 118 , wherein said one or more computer processors are further individually or collectively programmed to determine lengths of at least a subset of said plurality of homopolymer sequences based at least in part on clustering of said plurality of homopolymer sequences or said plurality of sequence signals associated with said plurality of homopolymer sequences.

Join the waitlist — get patent alerts

Track US2020303039A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.