Imputation-based sampling rate adjustment of parallel data streams
Abstract
Techniques for generating imputation-based, uniformly sampled parallel streams of time-series data are disclosed. A system divides into two subsets a dataset made up of multiple data streams. The data streams include interpolated data. The system trains one data correlation model using one subset of the data and applies the trained model to the other subset. The system replaces the interpolated values in the other subset with estimated values generated by the model. The system trains another data correlation model using the revised subset. The system applies the new model to the initial subset to generate estimated values for the initial subset. The system replaces the interpolated values in the initial subset with the estimated values. The system repeats the process of training data correlation models and revising previously-interpolated data points in the subsets of data until a predetermined iteration threshold is met.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A non-transitory computer readable medium comprising instructions which, when executed by one or more hardware processors, causes performance of operations comprising:
identifying a dataset comprising a plurality of data points; wherein each data point in the plurality of data points comprises a plurality of values; partitioning the dataset into at least a first subset of data points and a second subset of data points, wherein the first subset of data points comprises a first data point and the second subset of data points comprises a second data point; wherein a first plurality of values, corresponding to the first data point, comprises (a) a first set of sensor-detected values and (b) a first set of interpolated values; wherein a second plurality of values, corresponding to a second data point, comprises (a) a second set of sensor-detected values and (b) a second set of interpolated values; updating interpolated values in the first subset of data points and the second subset of data points at least by:
(a) training a first data correlation model based on the first subset of data points;
(b) applying the first data correlation model to at least a portion of the second subset of data points to revise one or more interpolated values of the second subset of data points to generate a revised second subset of data points;
(c) training a second data correlation model based on the revised second subset of data points;
(d) applying the second data correlation model to at least a portion of the first subset of data points to revise one or more interpolated values of the first subset of data points to generate a revised first subset of data points;
concatenating the revised first subset of data points and the revised second subset of data points to generate a revised plurality of data points.
2 . The medium of claim 1 , wherein applying the first data correlation model to at least the portion of the second subset of data points to revise the one or more interpolated values of the second subset of data points comprises revising at least one of the second set of interpolated values comprised in the second data point.
3 . The medium of claim 1 , wherein the operations further comprise:
(e) training a third data correlation model based on the revised first subset of data points; (b) applying the third data correlation model to at least a portion of the revised second subset of data points to revise the one or more interpolated values of the second subset of data points to further revise the revised second subset of data points.
4 . The medium of claim 1 , wherein the operations further comprise:
identifying a process utilizing the plurality of data points to perform one or more operations; and providing the revised plurality of data points to the process, instead of the plurality of data points, to perform the one or more operations.
5 . The medium of claim 1 , wherein the operations further comprise:
(e) subsequent to revising the one or more interpolated values of the first subset of data points and the one or more interpolated values of the second subset of data points, repeating operations (a)-(d), replacing the first subset with the revised first subset, and replacing the second subset with the revised second subset.
6 . The medium of claim 1 , wherein the dataset comprises a plurality of parallel time-series sensor data signals,
wherein partitioning the dataset into at least a first subset of data points and a second subset of data points comprises dividing the dataset into the first subset of data points generated prior to a particular time and the second subset of data points generated at, or after, the particular time.
7 . The medium of claim 1 , wherein the first data correlation model and the second data correlation model are multivariate state estimation technique (MSET) models.
8 . The medium of claim 1 , wherein the first set of interpolated values and the second set of interpolated values are generated by up-sampling data streams from one or more sensors.
9 . A method, comprising:
identifying a dataset comprising a plurality of data points; wherein each data point in the plurality of data points comprises a plurality of values; partitioning the dataset into at least a first subset of data points and a second subset of data points, wherein the first subset of data points comprises a first data point and the second subset of data points comprises a second data point; wherein a first plurality of values, corresponding to the first data point, comprises (a) a first set of sensor-detected values and (b) a first set of interpolated values; wherein a second plurality of values, corresponding to a second data point, comprises (a) a second set of sensor-detected values and (b) a second set of interpolated values; updating interpolated values in the first subset of data points and the second subset of data points at least by:
(a) training a first data correlation model based on the first subset of data points;
(b) applying the first data correlation model to at least a portion of the second subset of data points to revise one or more interpolated values of the second subset of data points to generate a revised second subset of data points;
(c) training a second data correlation model based on the revised second subset of data points;
(d) applying the second data correlation model to at least a portion of the first subset of data points to revise one or more interpolated values of the first subset of data points to generate a revised first subset of data points;
concatenating the revised first subset of data points and the revised second subset of data points to generate a revised plurality of data points.
10 . The method of claim 9 , wherein applying the first data correlation model to at least the portion of the second subset of data points to revise the one or more interpolated values of the second subset of data points comprises revising at least one of the second set of interpolated values comprised in the second data point.
11 . The method of claim 9 , further comprising:
(e) training a third data correlation model based on the revised first subset of data points; (b) applying the third data correlation model to at least a portion of the revised second subset of data points to revise the one or more interpolated values of the second subset of data points to further revise the revised second subset of data points.
12 . The method of claim 9 , further comprising:
identifying a process utilizing the plurality of data points to perform one or more operations; and providing the revised plurality of data points to the process, instead of the plurality of data points, to perform the one or more operations.
13 . The method of claim 9 , further comprising:
(e) subsequent to revising the one or more interpolated values of the first subset of data points and the one or more interpolated values of the second subset of data points, repeating operations (a)-(d), replacing the first subset with the revised first subset, and replacing the second subset with the revised second subset.
14 . The method of claim 9 , wherein the dataset comprises a plurality of parallel time-series sensor data signals,
wherein partitioning the dataset into at least a first subset of data points and a second subset of data points comprises dividing the dataset into the first subset of data points generated prior to a particular time and the second subset of data points generated at, or after, the particular time.
15 . The method of claim 9 , wherein the first data correlation model and the second data correlation model are multivariate state estimation technique (MSET) models.
16 . The method of claim 9 , wherein the first set of interpolated values and the second set of interpolated values are generated by up-sampling data streams from one or more sensors.
17 . A system, comprising:
one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the system to perform: identifying a dataset comprising a plurality of data points; wherein each data point in the plurality of data points comprises a plurality of values; partitioning the dataset into at least a first subset of data points and a second subset of data points, wherein the first subset of data points comprises a first data point and the second subset of data points comprises a second data point; wherein a first plurality of values, corresponding to the first data point, comprises (a) a first set of sensor-detected values and (b) a first set of interpolated values; wherein a second plurality of values, corresponding to a second data point, comprises (a) a second set of sensor-detected values and (b) a second set of interpolated values; updating interpolated values in the first subset of data points and the second subset of data points at least by:
(a) training a first data correlation model based on the first subset of data points;
(b) applying the first data correlation model to at least a portion of the second subset of data points to revise one or more interpolated values of the second subset of data points to generate a revised second subset of data points;
(c) training a second data correlation model based on the revised second subset of data points;
(d) applying the second data correlation model to at least a portion of the first subset of data points to revise one or more interpolated values of the first subset of data points to generate a revised first subset of data points;
concatenating the revised first subset of data points and the revised second subset of data points to generate a revised plurality of data points.
18 . The system of claim 17 , wherein applying the first data correlation model to at least the portion of the second subset of data points to revise the one or more interpolated values of the second subset of data points comprises revising at least one of the second set of interpolated values comprised in the second data point.
19 . The system of claim 17 , wherein the instructions cause the system to further perform:
(e) training a third data correlation model based on the revised first subset of data points; (b) applying the third data correlation model to at least a portion of the revised second subset of data points to revise the one or more interpolated values of the second subset of data points to further revise the revised second subset of data points.
20 . The system of claim 17 , wherein the instructions cause the system to further perform:
identifying a process utilizing the plurality of data points to perform one or more operations; and providing the revised plurality of data points to the process, instead of the plurality of data points, to perform the one or more operations.Join the waitlist — get patent alerts
Track US2022383033A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.