Methods, Systems, and Computer Readable Media for Evaluating Variant Likelihood
Abstract
A method for evaluating variant likelihood includes: providing a plurality of template polynucleotide strands, sequencing primers, and polymerase in a plurality of defined spaces disposed on a sensor array; exposing the plurality of template polynucleotide strands, sequencing primers, and polymerase to a series of flows of nucleotide species according to a predetermined order; obtaining measured values corresponding to an ensemble of sequencing reads for at least some of the template polynucleotide strands in at least one of the defined spaces; and evaluating a likelihood that a variant sequence is present given the measured values corresponding to the ensemble of sequencing reads, the evaluating comprising: determining a measurement confidence value for each read in the ensemble of sequencing reads and modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands.
Claims
exact text as granted — not AI-modified1 . A method for evaluating variant likelihood in nucleic acid sequencing, comprising:
(a) providing a plurality of template polynucleotide strands, sequencing primers, and polymerase in a plurality of defined spaces disposed on a sensor array; (b) exposing the plurality of template polynucleotide strands, sequencing primers, and polymerase to a series of flows of nucleotide species according to a predetermined order; (c) obtaining measured values corresponding to an ensemble of sequencing reads for at least some of the template polynucleotide strands in at least one of the defined spaces; and (d) evaluating a likelihood that a variant sequence is present given the measured values corresponding to the ensemble of sequencing reads, the evaluating comprising:
(i) determining a measurement confidence value for each read in the ensemble of sequencing reads, wherein the determining is based on variations between the measured values and model-predicted values for hypothesized sequences obtained using a predictive model of nucleotide incorporations responsive to flows of nucleotide species; and
(ii) modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands, wherein the modifying is based on variations between model-predicted values for different hypothesized sequences obtained using the predictive model of nucleotide incorporations responsive to flows of nucleotide species.
2 . The method of claim 1 , wherein modifying the at least some model-predicted values comprises applying a transformation including a product of (i) one of the first and second biases and (ii) a discriminant vector representing a difference between model-predicted values corresponding to different hypothesized sequences.
3 . The method of claim 1 , wherein evaluating the likelihood further comprises assigning a first frequency to a variant sequence and a second frequency to a non-variant sequence, and calculating a likelihood of having observed the ensemble of sequencing reads conditioned on the first frequency as a function of a product of the likelihoods of having observed each of the sequencing reads given the first frequency.
4 . The method of claim 1 , wherein evaluating the likelihood further comprises assigning a first frequency to a variant sequence, a second frequency to a non-variant sequence, and a third frequency to an outlier event.
5 . The method of claim 4 , wherein the outlier event has a flat density across all sequencing reads in the ensemble.
6 . The method of claim 4 , wherein evaluating the likelihood further comprises calculating a likelihood of having observed the ensemble of sequencing reads conditioned on the third frequency as a function of a product of the likelihoods of having observed each of the sequencing reads given the third frequency.
7 . The method of claim 1 , wherein the measurement confidence values are estimated using a function comprising a sum of log-likelihood of values measured for a given flow given a hypothesized sequence.
8 . The method of claim 7 , wherein the measurement confidence values are estimated using a function comprising differences between the measured values and the model-predicted values.
9 . The method of claim 8 , wherein the differences between measured and model-predicted values at each nucleotide flow are assumed to follow independent normal distributions each having a mean and a variance.
10 . The method of claim 9 , wherein the differences between measured and model-predicted values at each nucleotide flow are assumed to follow independent t-distributions.
11 . The method of claim 1 , wherein the measurement confidence values are estimated using an expression comprising
ε
i
=
1
1
+
exp
(
LL
yi
-
LL
xi
)
where LL yi and LL xi are log-likelihoods of values measured for a given sequencing read under hypothesized sequences y and x, respectively.
12 . The method of claim 11 , wherein the measurement confidence values are estimated using an expression for responsibility comprising
ρ
i
=
π
π
+
(
1
-
π
)
*
exp
(
LL
yi
-
LL
xi
)
where π represents a first frequency assigned to a variant sequence, 1−π represents a second frequency assigned to a non-variant sequence, and ρ i represents a measure of responsibility for each of the sequencing reads in the ensemble.
13 . The method of claim 11 , where the variance is estimated by decomposition of the variance in a flow and sequencing read into underlying latent components.
14 . The method of claim 13 , wherein each latent component corresponds to a homopolymer having an integer length.
15 . The method of claim 13 , wherein the latent components include a null variance component representing contribution to a flow regardless of any nucleotide incorporation, a residual variance component representing contribution for nucleotide incorporations not explicitly modeled, and one or more additional variance components.
16 . The method of claim 15 , wherein the one or more additional variance components comprise variance components associated with homopolymers having an integer length.
17 . The method of claim 13 , wherein the latent components are estimated using an EM methodology and a method of moments approximation.
18 . The method of claim 1 , wherein evaluating the likelihood comprises estimating (i) a first frequency π assigned to a variant sequence, (ii) at least one of a measurement confidence value ε i for each of the sequencing reads in the ensemble and a measure of responsibility ρ i for each of the sequencing reads in the ensemble, and (iii) a variance σ ij 2 for each of the flows and sequencing reads in the ensemble.
19 . A non-transitory machine-readable storage medium comprising instructions which, when executed by a processor, cause the processor to perform a method for evaluating variant likelihood in nucleic acid sequencing comprising:
(a) obtaining measured values corresponding to an ensemble of sequencing reads for at least some template polynucleotide strands in at least one defined space, wherein a plurality of template polynucleotide strands, sequencing primers, and polymerase have been provided in a plurality of defined spaces disposed on a sensor array, and wherein the plurality of template polynucleotide strands, sequencing primers, and polymerase have been exposed to a series of flows of nucleotide species according to a predetermined order; and (b) evaluating a likelihood that a variant sequence is present given the measured values corresponding to the ensemble of sequencing reads, the evaluating comprising:
(i) determining a measurement confidence value for each read in the ensemble of sequencing reads, wherein the determining is based on variations between the measured values and model-predicted values for hypothesized sequences obtained using a predictive model of nucleotide incorporations responsive to flows of nucleotide species; and
(ii) modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands, wherein the modifying is based on variations between model-predicted values for different hypothesized sequences obtained using the predictive model of nucleotide incorporations responsive to flows of nucleotide species.
20 . A system for evaluating variant likelihood in nucleic acid sequencing, including:
a plurality of template polynucleotide strands, sequencing primers, and polymerase provided in a plurality of defined spaces disposed on a sensor array; an apparatus configured to expose the plurality of template polynucleotide strands, sequencing primers, and polymerase to a series of flows of nucleotide species according to a predetermined order; a machine-readable memory; and a processor configured to execute machine-readable instructions, which, when executed by the processor, cause the system to perform a method for evaluating variant likelihood, comprising: (a) obtaining measured values corresponding to an ensemble of sequencing reads for at least some of the template polynucleotide strands in at least one of the defined spaces; and (b) evaluating a likelihood that a variant sequence is present given the measured values corresponding to the ensemble of sequencing reads, the evaluating comprising:
(i) determining a measurement confidence value for each read in the ensemble of sequencing reads, wherein the determining is based on variations between the measured values and model-predicted values for hypothesized sequences obtained using a predictive model of nucleotide incorporations responsive to flows of nucleotide species; and
(ii) modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands, wherein the modifying is based on variations between model-predicted values for different hypothesized sequences obtained using the predictive model of nucleotide incorporations responsive to flows of nucleotide species.Join the waitlist — get patent alerts
Track US2014296080A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.