Methods, systems, and computer readable media for evaluating variant likelihood
Abstract
A method comprises receiving an ensemble of sequencing reads based on measurements from a plurality of microwells of a sensor array; assigning measured values to the ensemble of sequencing reads; calculating model-predicted values utilizing a predictive model of nucleotide incorporations resulting from flows of nucleotide species according to a predetermined order; modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands, the modifying based on variations between model-predicted values for different hypothesized sequences obtained using the predictive model of nucleotide incorporations resulting from the flows of nucleotide species according to the predetermined order; calculating a measurement confidence value for each read in the ensemble of sequencing reads, the confidence value representing variations between the measured values and the modified model-predicted values; and identifying a plurality of reads in the ensemble as corresponding to a variant sequence.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of identifying sequence variants, the method comprising:
receiving an ensemble of sequencing reads based on measurements obtained from a plurality of microwells of a sensor array, the measurements indicative of incorporation of nucleotide species into a plurality of template polynucleotide strands disposed in the plurality of microwells of the sensor array; assigning measured values to the ensemble of sequencing reads; calculating model-predicted values for the plurality of template polynucleotide strands utilizing a predictive model of nucleotide incorporations resulting from flows of nucleotide species according to a predetermined order; modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands, wherein the modifying is based on variations between model-predicted values for different hypothesized sequences obtained using the predictive model of nucleotide incorporations resulting from the flows of nucleotide species according to the predetermined order; calculating a measurement confidence value for each read in the ensemble of sequencing reads, wherein the measurement confidence value represents variations between the measured values and the modified model-predicted values; and identifying a plurality of reads in the ensemble as corresponding to a variant sequence based on the measurement confidence value.
2 . The method of claim 1 , wherein the modifying comprises applying a transformation including a product of one of the first bias and the second bias and a discriminant vector representing a difference between vectors of model-predicted values corresponding to different hypothesized sequences.
3 . The method of claim 1 , further comprising:
assigning a first frequency to the variant sequence; assigning a second frequency to a non-variant sequence; and calculating a likelihood of observing each of the variant sequence and the non-variant sequence in the ensemble of sequencing reads based on the first and second frequencies.
4 . The method of claim 3 , further comprising assigning a third frequency to an outlier event, wherein the outlier event is an event in which a sequence other than the variant sequence and the non-variant sequence occurs.
5 . The method of claim 4 , wherein the outlier event has a flat density across all sequencing reads in the ensemble.
6 . The method of claim 4 , wherein the calculating the likelihood further comprises calculating the likelihood of observing an outlier event in the ensemble based further on the third frequency.
7 . The method of claim 1 , wherein the measurement confidence values are estimated using a function comprising a sum of a log of the likelihood of values measured for a given flow given a hypothesized sequence of the different hypothesized sequences.
8 . The method of claim 7 , wherein the measurement confidence values are calculated using a function comprising differences between the measured values and the model-predicted values.
9 . The method of claim 8 , wherein the differences between measured and model-predicted values at each nucleotide flow are modeled to follow independent normal distributions each having a mean and a variance.
10 . The method of claim 9 , wherein the differences between measured and model-predicted values at each nucleotide flow are modeled to follow independent t-distributions.
11 . The method of claim 9 , wherein the variance is estimated by decomposition of the variance in a flow and sequencing read into underlying latent components.
12 . The method of claim 11 , wherein each latent component corresponds to a homopolymer having an integer length.
13 . The method of claim 11 , wherein the latent components include a null variance component representing contribution to a flow regardless of any nucleotide incorporation, a residual variance component representing contribution for nucleotide incorporations not explicitly modeled, and one or more additional variance components.
14 . The method of claim 13 , wherein the one or more additional variance components comprise variance components associated with homopolymers having an integer length.
15 . The method of claim 11 , wherein the latent components are estimated using an expectation-maximization (EM) methodology and a method of moments approximation.
16 . The method of claim 1 , wherein the measurement confidence values are calculated using an expression comprising
∈ i = 1 1 + e x p L L y i − L L x i where LL yi and LL xi are log-likelihoods of values measured for a given sequencing read under hypothesized sequences y and x, respectively.
17 . The method of claim 16 , wherein the measurement confidence values are calculated using an expression for responsibility comprising
ρ i = π π + 1 − π ∗ e x p L L y i − L L x i where π represents a first frequency assigned to a variant sequence, 1 - π represents a second frequency assigned to a non-variant sequence, and ρ i represents a measure of responsibility for each of the sequencing reads in the ensemble.
18 . The method of claim 1 , further comprising estimating (i) a first frequency π assigned to a variant sequence, (ii) at least one of a measurement confidence value ∈ i for each of the sequencing reads in the ensemble and a measure of responsibility ρ i for each of the sequencing reads in the ensemble, and (iii) a variance
σ i j 2
for each of the flows and sequencing reads in the ensemble.
19 . A non-transitory machine-readable storage medium comprising instructions which, when executed by a processor, to cause the processor to perform a method comprising:
receiving an ensemble of sequencing reads based on measurements obtained from a plurality of microwells of a sensor array, the measurements indicative of incorporation of nucleotide species into a plurality of template polynucleotide strands disposed in the plurality of microwells of the sensor array; assigning measured values to the ensemble of sequencing reads; calculating model-predicted values for the plurality of template polynucleotide strands utilizing a predictive model of nucleotide incorporations resulting from flows of nucleotide species according to a predetermined order; modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands, wherein the modifying is based on variations between model-predicted values for different hypothesized sequences obtained using the predictive model of nucleotide incorporations resulting from the flows of nucleotide species according to the predetermined order; calculating a measurement confidence value for each read in the ensemble of sequencing reads, wherein the measurement confidence value represents variations between the measured values and the modified model-predicted values; and identifying a plurality of reads in the ensemble as corresponding to a variant sequence based on the measurement confidence value.
20 . A system for identifying sequence variants, including:
a microwell array comprising a plurality of microwells comprising a plurality of template polynucleotide strands, sequencing primers, and polymerase; a sensor array comprising a plurality of sensors aligned with the plurality of microwells; a machine-readable memory operably coupled to the sensor array; and a processor operably coupled to the sensor array and the machine-readable memory, the processor being configured to execute machine-readable instructions, which, when executed by the processor, causes the system to perform a method comprising:
receiving an ensemble of sequencing reads based on measurements obtained from the plurality of sensors of the sensor array, the measurements indicative of incorporation of nucleotide species into a plurality of template polynucleotide strands disposed in the plurality of microwells;
assigning measured values to the ensemble of sequencing reads;
calculating model-predicted values for the plurality of template polynucleotide strands utilizing a predictive model of nucleotide incorporations resulting from flows of nucleotide species according to a predetermined order;
modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands, wherein the modifying is based on variations between model-predicted values for different hypothesized sequences obtained using the predictive model of nucleotide incorporations resulting from the flows of nucleotide species according to the predetermined order;
calculating a measurement confidence value for each read in the ensemble of sequencing reads, wherein the measurement confidence value represents variations between the measured values and the modified model-predicted values; and
identifying a plurality of reads in the ensemble as corresponding to a variant sequence based on the measurement confidence value.Join the waitlist — get patent alerts
Track US2023360726A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.