Models for analyzing data from sequencing-by-synthesis operations
Abstract
Mathematical models for the analysis of signal data generated by sequencing of a polynucleotide strand using a pH-based method of detecting nucleotide incorporation(s). In an embodiment, the measured output signal from the reaction confinement region of a reactor array is mathematically modeled. The output signal may be modeled as a linear combination of one or more signal components, including a background signal component. This model is solved to determine the nucleotide incorporation signal. In another embodiment, the incorporation signal from the reaction confinement region of a reactor array is mathematically modeled.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A non-transitory machine-readable storage medium storing instructions which, when executed by a processor, cause the processor to perform a method comprising:
receiving signal data generated in response to chemical reactions resulting from a flow of a series of nucleotides onto a reactor array comprising a plurality of chemFET sensors and having multiple reaction confinement regions, one or more copies of the polynucleotide strand being located in a loaded reaction confinement region of the reactor array; processing the signal data from the loaded reaction confinement region by subtracting a background signal to obtain incorporation signal data; fitting an incorporation signal model comprising a system of non-linear differential equations to the incorporation signal data obtained from the processing, the system of non-linear differential equations comprising a first equation mathematically representing a rate of change in a nucleotide concentration in the loaded reaction confinement region and a second equation mathematically representing an amount of active polymerase in the loaded reaction confinement region; and determining, from the fit of the incorporation signal model, an estimate of a number of nucleotide incorporations resulting from the flowing of the nucleotide reagent.
3 . The non-transitory machine-readable storage medium of claim 2 , wherein the incorporation signal model comprises a third equation representing a flux of hydrogen ions generated in response to the chemical reactions between the one or more copies of the polynucleotide strand and the series of nucleotides.
4 . The non-transitory machine-readable storage medium of claim 2 , wherein fitting the incorporation signal model to the incorporation signal data comprises applying greater weight to a portion of the signal data that is earlier or later in time than another portion of the signal data.
5 . The non-transitory machine-readable storage medium of claim 2 , wherein the first equation mathematically representing a rate of change in the nucleotide concentration in the loaded reaction confinement region comprises a difference between a first term that is proportional to a nucleotide concentration gradient and a second term that represents a rate at which nucleotides are consumed by the chemical reactions in the loaded reaction confinement region.
6 . The non-transitory machine-readable storage medium of claim 2 , wherein the second equation mathematically representing the amount of active polymerase in the loaded reaction confinement region comprises a sum of numbers of polymerases [a n ] located at polynucleotide strand positions where n additional base incorporations are needed to complete a homopolymer of length M, where n varies from 1 to M.
7 . The non-transitory machine-readable storage medium of claim 2 , wherein the incorporation signal model further comprises a third equation for the rate of change in the amount of active polymerase in the loaded reaction confinement region.
8 . The non-transitory machine-readable storage medium of claim 7 , wherein the third equation mathematically representing the rate of change in the amount of active polymerase in the loaded reaction confinement region comprises a term including a product between a nucleotide concentration and a number of polymerases [a n ] located at a polynucleotide strand position where n additional base incorporations are needed to complete a homopolymer of length M.
9 . The non-transitory machine-readable storage medium of claim 2 , wherein the second equation mathematically representing the amount of active polymerase in the loaded reaction confinement region comprises a cumulative Poisson equation that calculates a probability that any given polynucleotide strand has not yet completed all base incorporations in a homopolymer of length M.
10 . The non-transitory machine-readable storage medium of claim 2 , wherein the mathematically representing the amount of active polymerase in the loaded confinement region comprises the following equation:
[
A
]
=
[
A
]
t
-
0
{
e
-
∫
0
t
k
[
dNTP
]
∑
i
=
0
M
-
1
(
∫
0
t
k
[
dNTP
]
)
i
i
!
}
wherein [A] a total number of active polymerase, [A] t=0 is a starting number of active polymerase in the loaded confinement region before any nucleotides have been incorporated, [dNTP] is a nucleotide concentration, M is a length of the polynucleotide in nucleotides, and k is a reaction rate coefficient for the polymerase.
11 . A method of sequencing a polynucleotide strand, the method comprising:
flowing a first nucleotide solution onto a reactor array comprising a plurality of chemFET sensors and multiple reaction confinement regions, one or more copies of the polynucleotide strand being located in a loaded confinement region of the multiple reaction confinement regions, wherein the first nucleotide solution comprises a first nucleotide known to be non-complementary for incorporation into the polynucleotide strand; receiving first output signal data from the loaded confinement region resulting from the flowing of the first nucleotide solution; determining a background signal for the loaded confinement region by fitting a model to the received first output signal data; flowing a second nucleotide solution onto the reactor array, wherein the second nucleotide solution comprises a second nucleotide resulting in incorporation of the second nucleotide to the polynucleotide strand; receiving a second output signal resulting from the flowing of the second nucleotide solution determining an incorporation signal from the loaded confinement region by fitting the model to the received second output signal data; and estimating a number of nucleotides incorporated into the polynucleotide strand, wherein the model comprises a function for a background signal component and a function for an incorporation signal component.
12 . The method of claim 11 , wherein the fitting the model to the received second output signal comprises applying greater weight to a portion of the second output signal data that is earlier or later in time than another portion of the second output signal data.
13 . The method of claim 11 , wherein the function for the background signal component comprises a rate parameter relating to a rate of change in an amount of hydrogen ions in the loaded reaction confinement region.
14 . The method of claim 13 , wherein the rate parameter is a product between a diffusion constant of the hydrogen ions in the loaded confinement region and a buffering capacity of the loaded confinement region.
15 . The method of claim 13 , wherein the function for the background signal component further comprises a ratio parameter relating to a ratio of change in the amount of hydrogen ions in a representative empty reaction confinement region relative to the rate of change in the amount of hydrogen ions in the loaded reaction confinement region.
16 . The method of claim 11 , wherein the function for the background signal component comprises an integral of the difference between the output signal of a representative empty reaction confinement region and the output signal of from the loaded reaction confinement region from one or more prior time frames.
17 . The method of claim 11 , wherein the function for the background signal component is derived from a flux of hydrogen ions between the between the first nucleotide solution of the first flow and the loaded reaction confinement region.
18 . The method of claim 11 , wherein fitting the model to the received first and second output signal data comprises using at least one of a regression analysis, a Bayesian technique, a least square analysis, and a Levenberg-Marquardt algorithm.
19 . The method of claim 11 , wherein estimating the number of nucleotides incorporated uses a maximum peak of the determined incorporation signal.
20 . The method of claim 11 , wherein estimating the number of nucleotides incorporated uses a comparison of the determined incorporation signal to a set of reference incorporation signal curves.Join the waitlist — get patent alerts
Track US2022383983A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.