Methods for normalization of experimental data
Abstract
Methods for normalization of experimental data with experiment-to-experiment variability. The experimental data may include biotechnology data or other data where experiment-to-experiment variability is introduced by an environment used to conduct multiple iterations of the same experiment. Deviations in the experimental data are measured between a central character and data values from multiple indexed data sets. The central character is a value of an ordered comparison determined from the multiple indexed data sets. The central character includes zero-order and low order central characters. Deviations between the central character and the multiple indexed data sets are removed by comparing the central character to the measured deviations from the multiple indexed data sets, thereby reducing deviations between the multiple indexed data sets and thus reducing experiment-to-experiment variability. Preferred embodiments of the present invention may be used to reduce intra-experiment and inter-experiment variability. When experiment-to-experiment variability is reduced or eliminated, comparison of experimental results can be used with a higher degree of confidence. Experiment-to-experiment variability is reduced for biotechnology data with new methods that can be used for bioinformatics or for other types of experimental data that are visual displayed (e.g., telecommunications data, electrical data for electrical devices, optical data, physical data, or other data). Experimental data can be consistently collected, processed and visually displayed with results that are accurate and not subject to experiment-to-experiment variability. Thus, intended experimental goals or results (e.g., determining polynucleotide sequences such as DNA, cDNA, or mRNA sequences) may be achieved in a more efficient and effective manner.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method for data normalization for a plurality of indexed data sets, comprising the following steps:
measuring deviations from a determined central character and data values from a plurality of indexed data sets, wherein the determined central character is a mode of an ordered comparison determined from the plurality of indexed data sets; and removing deviations between the determined central character and the plurality of indexed data sets by comparing the determined central character to the measured deviations from the plurality of indexed data sets, thereby reducing deviations between the plurality of indexed data sets.
2 . A computer readable medium having stored therein instructions for causing a central processing unit to execute the method of claim 1 .
3 . The method of claim 1 wherein the determined central character is determined by applying a transform to data values from the plurality of indexed data sets to utilize data information across indices from the plurality of indexed data sets.
4 . The method of claim 1 wherein the determined central character is determined by applying any of a zero-order transform or a low-order transform.
5 . The method of claim 4 wherein the zero-order transform includes applying a constant to transform data points in the plurality of indexed data sets, wherein the constant is independent of data values in the plurality of indexed data sets.
6 . The method of claim 4 wherein the low-order transform includes applying a smoothly varying scaling function to transform data points in the plurality of indexed data sets, wherein the varying scaling function is dependent on data values in the plurality of indexed data sets.
7 . The method of claim 1 wherein the plurality of indexed data sets include processed polynucleotide data suitable for visual display.
8 . The method of claim 7 wherein the polynucleotide data includes any of DNA, cDNA, or mRNA data.
9 . The method of claim 1 wherein the removing step includes removing deviations between the plurality of indexed data sets to reduce experiment-to-experiment variability and make the plurality of indexed data sets suitable for comparison.
10 . The method of claim 9 wherein the comparison includes a visual comparison on a display device.
11 . A method for creating a zero-order central character, comprising the following steps:
removing data points from outer quantiles of a plurality of indexed data sets with a smoothing window to create a plurality of smoothed sets of data points; determining a set of indexed data set ratios from the plurality of smoothed sets of data points, wherein the set indexed data set ratios is determined by comparing a selected smoothed set of data points from a selected index data set to other smoothed sets of data points from other indexed data sets from the plurality of indexed data sets; removing outer quantiles of ratios from the set of indexed data set ratios to create a subset of indexed data set ratios; and determining an averaged set of ratios from ratios in the subset of indexed data set ratios to create a zero-order central character.
12 . A computer readable medium having stored therein instructions for causing a central processing unit to execute the method of claim 11 .
13 . The method of claim 11 wherein the step of removing data points includes removing data points with:
f ** k ≡[2/( P+ 2)]Σ p=−[P/2, ]. . . ,[P/2] [( P+ 2 −|p| )/( P+ 2)] f * k+p , wherein f ** k is a smoothed set of data points, P is size of a smoothing window for a set of data points-p from a k th -indexed data set, and f * is a data envelope enclosing a set of data points-p that does not include data points from outer quantiles of the k th -indexed data set.
14 . The method of claim 11 wherein the step of determining a set of indexed data set ratios includes determining:
(g ** k /f ** k ), wherein f ** k is a selected smooth set of data points from a selected k th indexed data set, and g ** k is another smoothed set of data points other than f ** k .
15 . The method of claim 11 wherein the step of removing outer quantiles of ratios includes removing outer quantiles of ratios with:
r k ( g,f )≡{ g ** k /f ** k :D s ( f ** )≦ f ** k ≦D t ( f ** ); D s ( g ** )≦ g ** k ≦D t ( g ** )}, wherein r k (g,f) is an indexed data set of ratios between a selected smooth set of data points f ** k from k th -indexed data sets, g ** k is another smoothed set of data points other than f ** k , D s (f ** ) is a s-th quantile of values in the selected smooth set of data points f ** k , D t (f ** ) is a t-th quantile of values in another smooth set of data points f ** k , D s (g ** ) is a s-th quantile of values in selected smooth set of data points g ** k , and D t (g ** ) is a t-th quantile of values in the other smooth set of data points g ** k .
16 . The method of claim 11 wherein the step of determining an averaged ratio from ratios in the subset of indexed data set ratios includes determining:
λ 0 ( f )≡avg(∀ k,g≠f ){ r k ( g,f ): D u ( r ( g,f ))≦ r k ( g,f ) ≦ D v ( r ( g,f ))}, wherein λ 0 (f) is a zero order central character, avg is an average, r k (g,f) is a k- th indexed data set ratio between a selected smoothed set of data points-f and another smoothed set of data points-g, other than f, D u (r(g,f)) is a u-th quantile of ratios r(g,f), and D v (r(g,f)) is a v-th quantile of ratios r(g,f).
17 . A method for data normalization, comprising the following steps:
measuring deviations from a zero-order central character and a plurality of indexed data sets, wherein the zero-order central character is determined from plurality of indexed data sets; and removing deviations between the zero-order central character and the plurality of indexed data sets with ratios between the zero-order central character and the plurality of index data sets to and with ratios between the plurality of indexed data sets an averaged set of ratios for the plurality of indexed data sets.
18 . A computer readable medium having stored therein instructions for causing a central processing unit to execute the method of claim 17 .
19 . The method of claim 17 wherein the plurality of indexed data sets include processed polynucleotide data suitable for visual display.
20 . The method of claim 19 wherein the polynucleotide data includes any of DNA, cDNA, or mRNA data.
21 . The method of claim 19 wherein the removing step includes removing deviations between the plurality of indexed data sets with a zero-order central character to reduce experiment-to-experiment variability and make the plurality of indexed data set suitable for comparison.
22 . The method of claim 21 wherein the comparison includes a visual comparison on a display device.
23 . A method for creating a low-order central character, comprising the following steps:
removing data points from outer quantiles of a plurality of indexed data sets with a smoothing window to create a plurality of smoothed sets of data points for the plurality of indexed data sets; determining a set of indexed data set ratios from the plurality of smoothed sets of data points, wherein the set of indexed data set ratios is determined by comparing a selected smoothed set of data points from a selected indexed data set to other smoothed sets of data points from other indexed data sets from the plurality of indexed data sets; creating logarithms of the set of indexed data set ratios to create a set of logarithm ratios; filtering the set of logarithm ratios to create a filtered set of logarithm ratios; and applying an exponentiation to an average of the filtered set of logarithm ratios to create a low-order central character.
24 . A computer readable medium having stored therein instructions for causing a central processing unit to execute the method of claim 23 .
25 . The method of claim 23 wherein the step of removing data points includes removing data points with:
f ** k ≡[2 /( P+ 2)]Σ p=−[P/2], . . . ,[P/2] [( P+ 2−| p| )/( P+ 2)] f * k+p , wherein f ** k is a smoothed set of data points, P is size of a smoothing window for a set of data points-p from a k th -indexed data set, and f * is a data envelope enclosing a set of data points-p that does not include data points from outer quantiles of the k th -indexed data set.
26 . The method of claim 23 wherein the step of determining a set of indexed data set ratios includes determining:
(g ** k /f ** k ), wherein f ** k is a selected smoothed set of data points from a selected k th indexed data set, and g ** k is another smoothed set of data points other than f ** k .
27 . The method of claim 23 wherein the step of creating logarithms of the set of indexed data set ratios to create a set of logarithm ratios includes applying:
log x (g ** k /f ** k ), wherein log, is a logarithm for a desired base-x, f ** k is a selected smoothed set of data points from a selected k th -indexed set of data points, g ** k is another smoothed set of data points other than f ** k .
28 . The method of claim 23 wherein the step of filtering the set of logarithm ratios to create a filtered set of logarithm ratios includes applying:
ρ k(g,f) ≡χ ω [log x ( g ** k /f ** k )], wherein ρ k(g,f) is a filtered set of logarithm ratios, χ ω is a filter, log x is a logarithm for a desired base-x, f ** k is a selected smooth set of data points from a selected k th indexed set of data points, g ** k is another smoothed set of data points other than f ** k .
29 . The method of claim 28 wherein the filter χ ω is a low pass filter.
30 . The method of claim 23 wherein the step of applying an exponentiation to an average of the filtered set of logarithm ratios includes applying:
λ k ( f )≡exp x [avg(∀ k, g≠f ){ρ k ( g,f )}/2], wherein λ k (f) is a low-order central character, exp x is an exponential for a desired base-x, avg is an average, and {ρ k (g,f} is a filtered set of logarithm ratios for a k th indexed data set.
31 . A method for data normalization, comprising the following steps:
measuring deviations from a low-order central character and a plurality of indexed data sets, wherein the low-order central character is determined from plurality of indexed data sets; removing deviations between the low-order central character and the multiple indexed data sets with ratio s between the low-order central character and filtered logarithms of ratios for the multiple indexed data sets and with an exponential of the filtered logarithms of ratios.
32 . A computer readable medium having stored therein instructions for causing a central processing unit to execute the method of claim 31 .
33 . The method of claim 31 wherein the plurality of indexed data sets include processed polynucleotide data suitable for visual display.
34 . The method of claim 33 wherein the polynucleotide data includes any of DNA, cDNA, or mRNA data.
35 . The method of claim 31 wherein the removing step includes removing deviations between the plurality of indexed data sets with a low-order central character to reduce experiment-to-experiment variability and make the plurality of indexed data sets suitable for comparison.
36 . The method of claim 35 wherein the comparison includes a visual comparison on a display device.
37 . A method for data normalization, comprising the following steps:
reading a plurality of indexed data sets, wherein the plurality of indexed data sets were produced by completing a desired experiment a plurality of times and wherein the plurality of indexed data sets include deviations in results for the desired experiment due to environment conditions used to complete the desired experiment a plurality of times; creating a central character from the plurality of indexed data sets; removing deviations between the central character and the plurality of indexed data sets by comparing the central character to measured deviations from the plurality of indexed data sets to create a normalized set of indexed data sets, thereby reducing experiment-to-experiment deviations among the plurality of indexed data sets for the desired experiment; and displaying the normalized set of indexed data sets on a display device for comparative analysis.
38 . A computer readable medium having stored therein instructions for causing a central processing unit to execute the method of claim 37 .
39 . The method of claim 37 wherein the plurality of indexed data sets include processed polynucleotide data suitable for visual display.
40 . The method of claim 39 wherein the polynucleotide data includes any of DNA, cDNA, or mRNA data.
41 . The method of claim 37 wherein the deviations due to environment conditions include deviations due to any of deviations in an electrophoresis gel or micro-arrays used to complete the desired experiment a plurality of times.
42 . The method of claim 37 wherein the central character is any of a zero-order central character or a low-order central character.
43 . The method of claim 37 wherein the step of creating a central character further comprises applying a normalization transform to data values from the plurality of indexed data sets to utilize data information across indices from the plurality of indexed data sets.
44 . The method of claim 43 wherein the normalization transform includes any of a zero-order transform or a low-order transform.Join the waitlist — get patent alerts
Track US2002049570A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.