Analysis of a polymer comprising polymer units
Abstract
A sequence of polymer units in a polymer ( 3 ), eg. DNA, is estimated from at least one series of measurements related to the polymer, eg. ion current as a function of translocation through a nanopore ( 1 ), wherein the value of each measurement is dependent on a k-mer being a group of k polymer units ( 4 ). A probabilistic model, especially a hidden Markov model (HMM), is provided, comprising, for a set of possible k-mers: transition weightings representing the chances of transitions from origin k-mers to destination k-mers; and emission weightings in respect of each k-mer that represent the chances of observing given values of measurements for that k-mer. The series of measurements is analysed using an analytical technique, eg. Viterbi decoding, that refers to the model and estimates at least one estimated sequence of polymer units in the polymer based on the likelihood predicted by the model of the series of measurements being produced by sequences of polymer units. In a further embodiment, different voltages are applied across the nanopore during translocation in order to improve the resolution of polymer units.
Claims
exact text as granted — not AI-modified1 . A method of estimating a sequence of polymer units in a polymer from at least one series of measurements related to the polymer, wherein the value of each measurement is dependent on a k-mer, being a group of k polymer units where k is a positive integer, the method comprising:
providing a model comprising, for a set of possible k-mers:
transition weightings representing the chances of transitions from origin k-mers to destination k-mers; and
emission weightings in respect of each k-mer that represent the chances of observing given values of measurements for that k-mer; and
analysing the series of measurements using an analytical technique that refers to the model and estimating at least one estimated sequence of polymer units in the polymer based on the likelihood predicted by the model of the series of measurements being produced by sequences of polymer units.
2 . A method according to claim 1 , wherein at least one of the transition weightings and the emission weightings comprise values of non-binary variables.
3 . A method according to claim 2 , wherein both of the transition weightings and the emission weightings comprise values of non-binary variables.
4 . A method according to claim 1 , wherein the emission weightings represent non-zero chances of observing all possible measurements.
5 . (canceled)
6 . (canceled)
7 . A method according to claim 1 , wherein k is a plural integer.
8 . A method according to claim 7 , wherein the transition weightings represent non-zero chances of preferred transitions, being transitions from origin k-mers to destination k-mers that have a sequence in which the first (k-1) polymer units are the final (k-1) polymer units of the origin k-mer, and represent lower chances of non-preferred transitions, being transitions from origin k-mers to destination k-mers that have a sequence different from the origin k-mer and in which the first (k-1) polymer units are not the final (k-1) polymer units of the origin k-mer.
9 . A method according to claim 8 , wherein the transition weightings represent non-zero chances of at least some of said non-preferred transitions.
10 . A method according to claim 9 , wherein the transition weightings represent non-zero chances of non-preferred transitions from origin k-mers to destination k-mers that have a sequence wherein the first (k-2) polymer units are the final (k-2) polymer units of the origin k-mer.
11 . A method according to claim 1 , wherein the analytical technique is a probabilistic technique.
12 . A method according to claim 1 , wherein the transition weightings are probabilities, and/or the emission weightings are probabilities.
13 . A method according to claim 1 , wherein the model is a Hidden Markov Model.
14 . A method according to claim 1 , wherein the step of analysing further comprises deriving a quality score in respect of the or each estimated sequence that represents the likelihood predicted by the model of the series of measurements being produced by the estimated sequence of polymer units.
15 . A method according to claim 1 , wherein the step of analysing further comprises deriving quality scores in respect of individual k-mers corresponding to the estimated sequence of polymer units, that represent the likelihoods predicted by the model of the series of measurements being produced by a sequence including the individual k-mers and/or deriving quality scores in respect of sequences of k-mers corresponding to the estimated sequence of polymer units, that represent the likelihoods predicted by the model of the series of measurements being produced by the given sequences of k-mers.
16 . (canceled)
17 . A method according to claim 1 , wherein the step of analysing derives plural estimated sequences of polymer units in the polymer.
18 . (canceled)
19 . A method according to claim 1 , wherein the step of estimating at least one estimated sequence of polymer units in the polymer comprises:
estimating at least one sequence of k-mers based on the likelihood predicted by the model of the series of measurements being produced by overall sequences of k-mers; and estimating a sequence of polymer units from the estimated sequence of k-mers.
20 . A method according to claim 1 , wherein, in the at least one series of measurements, a predetermined number of measurements are dependent on each k-mer, the predetermined number being one or more.
21 . A method according to claim 20 , wherein
the method comprises receiving at least one input signal comprising an input series of measurements in which groups of plural measurements are dependent on the same k-mer, without a priori knowledge of the number of measurements in the group, and before the step of analysing, processing the at least one input signal to identify successive groups of measurements and to derive said predetermined number of measurements in respect of each identified group, the step of analysing being performed on the or each series of measurements thus derived.
22 . A method according to claim 1 , wherein, in the at least one series of measurements, groups of plural measurements are dependent on the same k-mer, without a priori knowledge of number of measurements in the group.
23 . (canceled)
24 . A method according to claim 1 , wherein said measurements of the polymer are made during translocation of the polymer through a nanopore
25 . (canceled)
26 . A method according to claim 24 wherein translocation of the polymer through the nanopore is performed in a ratcheted manner.
27 . A method according to claim 24 , wherein the polymer is a polynucleotide, and the polymer units are nucleotides.
28 . (canceled)
29 . A method according to claim 24 , wherein the nanopore is a biological pore.
30 . (canceled)
31 . A method according to claim 1 , wherein
the method is performed on plural series of measurements each related to said polymer, wherein the value of each measurement is dependent on a k-mer, and the analytical technique treats the plural series of measurements as arranged in plural, respective dimensions.
32 . A method according to claim 31 , wherein each series of measurements are measurements of the same region of the same polymer.
33 . A method according to claim 31 , wherein the plural series of measurements comprise two series of measurements, wherein the first series of measurements are measurements of a first region of a polymer and the second series of measurements are measurements of a second region of a polymer that is related to said first region.
34 . A method according to claim 33 , wherein the first and second regions are related regions of the same polymer.
35 . A method according to claim 33 , wherein the related regions are complementary.
36 . (canceled)
37 . (canceled)
38 . (canceled)
39 . An analysis device for estimating a sequence of polymer units in a polymer from at least one series of measurements related to the polymer, wherein the value of each measurement is dependent on a k-mer being a group of k polymer units, where k is a plural integer, the method comprising:
a memory storing a model comprising, for a set of possible k-mers:
transition weightings representing the chances of transitions from origin k-mers to destination k-mers; and
emission weightings in respect of each k-mer that represent the chances of observing given values of measurements for that k-mer; and
an analysis unit configured to analyse the series of measurements using an analytical technique that refers to the model and to estimate at least one estimated sequence of polymer units in the polymer based on the likelihood predicted by the model of the series of measurements being produced by sequences of polymer units.
40 . (canceled)
41 . (canceled)
42 . (canceled)
43 . (canceled)
44 . (canceled)
45 . (canceled)
46 . (canceled)
47 . (canceled)
48 . (canceled)
49 . (canceled)
50 . (canceled)
51 . (canceled)
52 . (canceled)
53 . (canceled)
54 . (canceled)
55 . (canceled)
56 . (canceled)
57 . (canceled)
58 . (canceled)
59 . (canceled)
60 . (canceled)
61 . (canceled)
62 . (canceled)
63 . (canceled)
64 . (canceled)
65 . (canceled)
66 . (canceled)
67 . (canceled)
68 . (canceled)
69 . (canceled)
70 . (canceled)Join the waitlist — get patent alerts
Track US2016162634A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.