Automatic selection of high quality training data using an adaptive oracle-trained learning framework
Abstract
In general, embodiments of the present invention provide systems, methods and computer readable media for an adaptive oracle-trained learning framework for automatically building and maintaining models that are developed using machine learning algorithms. In embodiments, the framework leverages at least one oracle (e.g., a crowd) for automatic generation of high-quality training data to use in deriving a model. Once a model is trained, the framework monitors the performance of the model and, in embodiments, leverages active learning and the oracle to generate feedback about the changing data for modifying training data sets while maintaining data quality to enable incremental adaptation of the model.
Claims
exact text as granted — not AI-modified1 - 32 . (canceled)
33 . A system, comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to:
calculate a set of confidence values for a set of features associated with multi-dimensional data, wherein each confidence value of the set of confidence values is associated with an operator estimate, wherein each confidence value represents a probability that a feature representation of the multi-dimensional data belongs to a distribution, and wherein the operator estimate is associated with an operator determined by a statistical model; update at least one training data set with the multi-dimensional data in response to a determination, based on the set of confidence values, that the multi-dimensional data is included in the at least one training data set; and generate at least one model from the at least one training data set.
34 . The system of claim 33 , wherein each feature from the set of features represents a value of a corresponding attribute of the multi-dimensional data.
35 . The system of claim 33 , wherein the operator estimate is configured to modify a feature from the set of features based on the statistical model.
36 . The system of claim 33 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
calculate an operator estimation score for the multi-dimensional data based on the feature representation of the multi-dimensional data.
37 . The system of claim 33 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
calculate an operator estimation score for the multi-dimensional data based on an estimator trained using the set of confidence values.
38 . The system of claim 33 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
calculate a set of global estimate confidence values, wherein global estimate confidence value from the set of global estimate confidence values represents a probability of the feature representation belonging to a global distribution represented by a global data set that contains one or more data instances of a same type as the multi-dimensional data.
39 . The system of claim 38 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
calculate a global estimation score based on the set of global estimate confidence values.
40 . The system of claim 33 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
calculate an estimation score for the multi-dimensional data based on the set of confidence values; and assign the multi-dimensional data to a class associated with the at least one model based on the estimation score for the multi-dimensional data.
41 . The system of claim 33 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
determine accuracy of the at least one model based on a confidence value associated with output of the at least one model; and update the at least one training data set based on the confidence value associated with the output of the at least one model.
42 . A computer-implemented method, comprising:
calculating, by a processor, a set of confidence values for a set of features associated with multi-dimensional data, wherein each confidence value of the set of confidence values is associated with an operator estimate, wherein each confidence value represents a probability that a feature representation of the multi-dimensional data belongs to a distribution, and wherein the operator estimate is associated with an operator determined by a statistical model; updating, by the processor, at least one training data set with the multi-dimensional data in response to a determination, based on the set of confidence values, that the multi-dimensional data is included in the at least one training data set; and generating, by the processor, at least one model from the at least one training data set.
43 . The computer-implemented method of claim 42 , further comprising:
calculating, by the processor, an operator estimation score for the multi-dimensional data based on the feature representation of the multi-dimensional data.
44 . The computer-implemented method of claim 42 , further comprising:
calculating, by the processor, an operator estimation score for the multi-dimensional data based on an estimator trained using the set of confidence values.
45 . The computer-implemented method of claim 42 , further comprising:
calculating, by the processor, a set of global estimate confidence values, wherein global estimate confidence value from the set of global estimate confidence values represents a probability of the feature representation belonging to a global distribution represented by a global data set that contains one or more data instances of a same type as the multi-dimensional data.
46 . The computer-implemented method of claim 45 , further comprising:
calculating, by the processor, a global estimation score based on the set of global estimate confidence values.
47 . The computer-implemented method of claim 42 , further comprising:
calculating, by the processor, an estimation score for the multi-dimensional data based on the set of confidence values; and assigning, by the processor, the multi-dimensional data to a class associated with the at least one model based on the estimation score for the multi-dimensional data.
48 . The computer-implemented method of claim 42 , further comprising:
determining, by the processor, accuracy of the at least one model based on a confidence value associated with output of the at least one model; and updating, by the processor, the at least one training data set based on the confidence value associated with the output of the at least one model.
49 . A computer program product, stored on a computer readable medium,
comprising instructions that when executed by one or more computers cause the one or more computers to: calculate a set of confidence values for a set of features associated with multi-dimensional data, wherein each confidence value from the set of confidence values is associated with an operator estimate, wherein each confidence value represents a probability that a feature representation of the multi-dimensional data belongs to a distribution, and wherein the operator estimate is associated with an operator determined by a statistical model; update at least one training data set with the multi-dimensional data in response to a determination, based on the set of confidence values, that the multi-dimensional data is included in the at least one training data set; and generate at least one model from the at least one training data set.
50 . The computer program product of claim 49 , wherein the instructions, when executed by the one or more computers, further cause the one or more computers to:
calculate an operator estimation score for the multi-dimensional data based on an estimator trained using the set of confidence values.
51 . The computer program product of claim 49 , wherein the instructions, when executed by the one or more computers, further cause the one or more computers to:
calculate a set of global estimate confidence values, wherein global estimate confidence value from the set of global estimate confidence values represents a probability of the feature representation belonging to a global distribution represented by a global data set that contains one or more data instances of a same type as the multi-dimensional data.
52 . The computer program product of claim 51 , wherein the instructions, when executed by the one or more computers, further cause the one or more computers to:
calculate a global estimation score based on the set of global estimate confidence values.Join the waitlist — get patent alerts
Track US2020302337A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.