US2020293951A1PendingUtilityA1
Dynamically optimizing a data set distribution
Est. expiryAug 19, 2034(~8.1 yrs left)· nominal 20-yr term from priority
G06N 7/01G06F 16/285G06F 16/2453G06F 16/906G06N 20/00G06F 16/23G06N 5/02
61
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
In general, embodiments of the present invention provide systems, methods and computer readable media configured to receive configuration data describing a desired data set distribution, and, in response to receiving new data instances, use the configuration data and the new data instances to dynamically optimize the distribution of data already stored in a data reservoir that has been discretized into bins representing the desired data distribution.
Claims
exact text as granted — not AI-modified1 - 42 . (canceled)
43 . A system, comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to:
receive a data set optimization job that comprises a set of input data associated with an input range and distribution configuration data that describes a discretization of a data set distribution of the set of input data into a plurality of data bins, wherein the discretization defines a range of evaluation values for a data bin from the plurality of data bins, and wherein the discretization further defines a label for the data bin; generate an evaluation determination that indicates whether to provide a portion of the input data to the data bin; in response to a determination that the evaluation determination satisfies a defined criterion, generate an updated determination for the data bin based on the range of evaluation values associated with the data bin and an instance evaluation of the portion of the input data, wherein the updated determination provides an indication to associate the portion of the input data instance with the data bin; and update the label for the data bin based on the updated determination for the data bin.
44 . The system of claim 43 , wherein the discretization further defines a prediction confidence score for the data bin to store the portion of the input data.
45 . The system of claim 43 , wherein the discretization further defines a size capacity for the data bin.
46 . The system of claim 43 , wherein the label for the bin identifies another portion of the input data that is associated with the data bin.
47 . The system of claim 43 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
determine whether an evaluation value of the portion of the input data is within the range of evaluation values associated with the data bin.
48 . The system of claim 43 , wherein the evaluation determination provides an indication as to whether a bin size capacity for the data bin is satisfied.
49 . The system of claim 43 , wherein the data set optimization job is associated with an input data evaluator.
50 . The system of claim 43 , wherein the data set optimization job is associated with a predictive model, and wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
generate a model prediction for the portion of the input data based on the predictive model.
51 . The system of claim 50 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
determine whether the model prediction for the portion of the input data is within the range of evaluation values associated with the data bin.
52 . The system of claim 43 , wherein the instance evaluation for the first input data instance is an anomaly score generated based on an anomaly scorer.
53 . The system of claim 43 , wherein the set of input data is received from a data stream.
54 . The system of claim 43 , wherein the set of input data is received from a data stream generated from online data.
55 . A computer-implemented method, comprising:
receiving a data set optimization job that comprises a set of input data associated with an input range and distribution configuration data that describes a discretization of a data set distribution of the set of input data into a plurality of data bins, wherein the discretization defines a range of evaluation values for a data bin from the plurality of data bins, and wherein the discretization further defines a label for the data bin; generating an evaluation determination that indicates whether to provide a portion of the input data to the data bin; in response determining that the evaluation determination satisfies a defined criterion, generating an updated determination for the data bin based on the range of evaluation values associated with the data bin and an instance evaluation of the portion of the input data, wherein the updated determination provides an indication to associate the portion of the input data instance with the data bin; and updating the label for the data bin based on the updated determination for the data bin.
56 . The computer-implemented method of claim 55 , further comprising:
defining a prediction confidence score for the data bin to store the portion of the input data.
57 . The computer-implemented method of claim 55 , further comprising:
determining whether an evaluation value of the portion of the input data is within the range of evaluation values associated with the data bin.
58 . The computer-implemented method of claim 55 , further comprising:
generating a model prediction for the portion of the input data based on a predictive model associated with the data set optimization job.
59 . The computer-implemented method of claim 58 , further comprising:
determining whether the model prediction for the portion of the input data is within the range of evaluation values associated with the data bin.
60 . A computer program product, stored on a computer readable medium, comprising instructions that when executed by one or more computers cause the one or more computers to:
receive a data set optimization job that comprises a set of input data associated with an input range and distribution configuration data that describes a discretization of a data set distribution of the set of input data into a plurality of data bins, wherein the discretization defines a range of evaluation values for a data bin from the plurality of data bins, and wherein the discretization further defines a label for the data bin; generate an evaluation determination that indicates whether to provide a portion of the input data to the data bin; in response to a determination that the evaluation determination satisfies a defined criterion, generate an updated determination for the data bin based on the range of evaluation values associated with the data bin and an instance evaluation of the portion of the input data, wherein the updated determination provides an indication to associate the portion of the input data instance with the data bin; and update the label for the data bin based on the updated determination for the data bin.
61 . The computer program product of claim 60 , wherein the discretization further defines a prediction confidence score for the data bin to store the portion of the input data.
62 . The computer program product of claim 60 , wherein the discretization further defines a size capacity for the data bin.Join the waitlist — get patent alerts
Track US2020293951A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.