US2013217006A1PendingUtilityA1
Algorithms for sequence determination
Assignee: PACIFIC BIOSCIENCES CALIFORNIAPriority: Nov 20, 2008Filed: Dec 31, 2012Published: Aug 22, 2013
Est. expiryNov 20, 2028(~2.3 yrs left)· nominal 20-yr term from priority
G16B 30/20G16B 30/10G16B 30/00G06F 19/22
62
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The present invention is generally directed to powerful and flexible methods and systems for consensus sequence determination from replicate biomolecule sequence data. It is an object of the present invention to improve the accuracy of consensus biomolecule sequence determination from replicate sequence data by providing methods for assimilating replicate sequence into a final consensus sequence more accurately than any one-pass sequence analysis system.
Claims
exact text as granted — not AI-modified1 - 46 . (canceled)
47 . A method for determining a consensus base call at a given location i in a nucleic acid molecule, the method comprising:
a) computing a plurality of alignment lattices, wherein each respective alignment lattice in the plurality of alignment lattices is an alignment lattice between (i) a respective read in a plurality of reads of the nucleic acid molecule and (ii) a respective potential template sequence in a plurality of potential template nucleic acid sequences for the nucleic acid molecule, wherein the plurality of potential template sequences comprises a plurality of unique template sequences for the nucleic acid molecule and wherein each potential template nucleic acid sequence in the plurality of potential nucleic acid sequences includes a position j that corresponding to location i in the nucleic acid molecule, thereby generating a corresponding set of alignment lattices for each respective potential template sequence in the plurality of potential template sequences; b) computing a likelihood for each respective alignment lattice in the plurality of alignment lattices using an error model for the plurality of reads; c) identifying a consensus template sequence from the plurality of potential template sequences based upon a likelihood computed for each respective alignment lattice in the plurality of alignment lattices; and d) deeming an identity of the base at location j in the consensus template sequence to be the consensus base call at the given location i in the nucleic acid molecule, wherein the computing a), computing b), and identifying c) are performed on a suitably programmed computer.
48 . The method of claim 47 , the method further comprising obtaining a read in the plurality of reads using a single-molecule sequencing technology.
49 . The method of claim 48 , wherein the single-molecule sequencing technology is selected from the group consisting of sequencing-by-synthesis and nanopore sequencing.
50 . The method of claim 47 , wherein the consensus base call is a substitution base call.
51 . The method of claim 47 , wherein the computing the plurality of alignment lattices comprises use of a CRF aligner.
52 . The method of claim 47 , wherein the computing the likelihood b) is performed using a recursive relation.
53 . The method of claim 52 , wherein the recursive relation computes both the likelihood summation over a plurality of paths ending at location (i,j) in an alignment lattice in the plurality of alignment lattices and a likelihood summation over all paths beginning at location (i,j) in the alignment lattice.
54 . The method of claim 47 , wherein the likelihood for each respective lattice is based on a plurality of signal features.
55 . The method of claim 47 , wherein the plurality of reads comprise both base calls and at least one of a (i) base call quality metric, (ii) a trace feature for a platform limitation metric, (iii) an expected error metric, and (iv) a detail of sequencing technology used to generate the plurality of sequence reads.
56 . The method of claim 47 , wherein the computing the likelihood for each respective alignment lattice in the plurality of alignment lattices uses an error model associated with the plurality of reads.
57 . A method for determining a consensus base call at a given location i in a nucleic acid molecule, the method comprising:
a) computing a plurality of alignment lattices, wherein each respective alignment lattice in the plurality of alignment lattices is an alignment lattice between (i) a respective read in a plurality of reads of the nucleic acid molecule and (ii) a template sequence, wherein the potential template nucleic acid sequence includes a position j that corresponds to location i in the nucleic acid molecule; b) identifying the base in the set {A, C, G, and T} for location i that maximizes a summation of a joint probability function summed across the plurality of reads of the nucleic acid molecule, the joint probability function comprising, for each respective read in the plurality of reads a product of (A) a forward path probability P given (i) the respective read, (ii) the template sequence, and (iii) an error model H associated with the plurality of reads, and (B) a backward path probability Q given (i) the respective read, (ii) the template sequence, and (iii) the error model H; and c) deeming the base identified by the identifying b) to be the consensus base call at the given location i in the nucleic acid molecule, wherein the computing a) and identifying b) are performed on a suitably programmed computer.
58 . The method of claim 47 , wherein the joint probability function is expressed as:
∑
n
∑
i
P
(
r
n
,
i
,
j
,
t
i
,
j
=
b
|
H
)
Q
(
r
1
,
i
,
j
,
t
i
,
j
=
b
|
H
)
,
wherein,
P(r n ,i,j,t i,j =b|H) is the forward path probability given the error model H,
Q(r n ,i,j,t i,j =b|H) is the backward path probability given the error model H,
n is a positive integer index to the plurality of reads,
r n is a read n in the plurality of reads, and
t i,j is the template at position j.
59 . The method of claim 58 , the method further comprising obtaining a read in the plurality of reads using a single-molecule sequencing technology.
60 . The method of claim 59 , wherein the single-molecule sequencing technology is selected from the group consisting of sequencing-by-synthesis and nanopore sequencing.
61 . The method of claim 58 , wherein the consensus base call is a substitution base call.
62 . The method of claim 58 , wherein the computing the plurality of alignment lattices comprises use of a CRF aligner.
63 . The method of claim 57 , wherein the joint probability function is computed using a recursive relation.
64 . An apparatus, comprising:
one or more processors; memory; and one or more programs stored in the memory for execution by the one or more processors, the one or more programs comprising instructions for: a) computing a plurality of alignment lattices, wherein each respective alignment lattice in the plurality of alignment lattices is an alignment lattice between (i) a respective read in a plurality of reads of the nucleic acid molecule and (ii) a respective potential template sequence in a plurality of potential template nucleic acid sequences for the nucleic acid molecule, wherein the plurality of potential template sequences comprises a plurality of unique template sequences for the nucleic acid molecule and wherein each potential template nucleic acid sequence in the plurality of potential nucleic acid sequences includes a position j that corresponding to location i in the nucleic acid molecule, thereby generating a corresponding set of alignment lattices for each respective potential template sequence in the plurality of potential template sequences; b) computing a likelihood for each respective alignment lattice in the plurality of alignment lattices using an error model for the plurality of reads; c) identifying a consensus template sequence from the plurality of potential template sequences based upon a likelihood computed for each respective alignment lattice in the plurality of alignment lattices; and d) deeming an identity of the base at location j in the consensus template sequence to be the consensus base call at the given location i in the nucleic acid molecule, wherein the computing a), computing b), and identifying c) are performed on a suitably programmed computer.
65 . An apparatus, comprising:
one or more processors; memory; and one or more programs stored in the memory for execution by the one or more processors, the one or more programs comprising instructions for: a) computing a plurality of alignment lattices, wherein each respective alignment lattice in the plurality of alignment lattices is an alignment lattice between (i) a respective read in a plurality of reads of the nucleic acid molecule and (ii) a template sequence, wherein the potential template nucleic acid sequence includes a position j that corresponds to location i in the nucleic acid molecule; b) identifying the base in the set {A, C, G, and T} for location i that maximizes a summation of a joint probability function summed across the plurality of reads of the nucleic acid molecule, the joint probability function comprising, for each respective read in the plurality of reads a product of (A) a forward path probability P given (i) the respective read, (ii) the template sequence, and (iii) an error model H associated with the plurality of reads, and (B) a backward path probability Q given (i) the respective read, (ii) the template sequence, and (iii) the error model H; and c) deeming the base identified by the identifying b) to be the consensus base call at the given location i in the nucleic acid molecule, wherein the computing a) and identifying b) are performed on a suitably programmed computer.
66 . The method of claim 47 , wherein the identifying comprises summing the likelihood computed for each respective alignment lattice in the plurality of alignment lattices.Join the waitlist — get patent alerts
Track US2013217006A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.