US2020302337A1PendingUtilityA1

Automatic selection of high quality training data using an adaptive oracle-trained learning framework

Assignee: GROUPON INCPriority: Dec 23, 2013Filed: Mar 4, 2020Published: Sep 24, 2020
Est. expiryDec 23, 2033(~7.4 yrs left)· nominal 20-yr term from priority
G06N 20/00
60
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In general, embodiments of the present invention provide systems, methods and computer readable media for an adaptive oracle-trained learning framework for automatically building and maintaining models that are developed using machine learning algorithms. In embodiments, the framework leverages at least one oracle (e.g., a crowd) for automatic generation of high-quality training data to use in deriving a model. Once a model is trained, the framework monitors the performance of the model and, in embodiments, leverages active learning and the oracle to generate feedback about the changing data for modifying training data sets while maintaining data quality to enable incremental adaptation of the model.

Claims

exact text as granted — not AI-modified
1 - 32 . (canceled) 
     
     
         33 . A system, comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to:
 calculate a set of confidence values for a set of features associated with multi-dimensional data, wherein each confidence value of the set of confidence values is associated with an operator estimate, wherein each confidence value represents a probability that a feature representation of the multi-dimensional data belongs to a distribution, and wherein the operator estimate is associated with an operator determined by a statistical model;   update at least one training data set with the multi-dimensional data in response to a determination, based on the set of confidence values, that the multi-dimensional data is included in the at least one training data set; and   generate at least one model from the at least one training data set.   
     
     
         34 . The system of  claim 33 , wherein each feature from the set of features represents a value of a corresponding attribute of the multi-dimensional data. 
     
     
         35 . The system of  claim 33 , wherein the operator estimate is configured to modify a feature from the set of features based on the statistical model. 
     
     
         36 . The system of  claim 33 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
 calculate an operator estimation score for the multi-dimensional data based on the feature representation of the multi-dimensional data.   
     
     
         37 . The system of  claim 33 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
 calculate an operator estimation score for the multi-dimensional data based on an estimator trained using the set of confidence values.   
     
     
         38 . The system of  claim 33 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
 calculate a set of global estimate confidence values, wherein global estimate confidence value from the set of global estimate confidence values represents a probability of the feature representation belonging to a global distribution represented by a global data set that contains one or more data instances of a same type as the multi-dimensional data.   
     
     
         39 . The system of  claim 38 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
 calculate a global estimation score based on the set of global estimate confidence values.   
     
     
         40 . The system of  claim 33 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
 calculate an estimation score for the multi-dimensional data based on the set of confidence values; and   assign the multi-dimensional data to a class associated with the at least one model based on the estimation score for the multi-dimensional data.   
     
     
         41 . The system of  claim 33 , wherein the one or more storage devices store instructions that are operable, when executed by the one or more computers, to further cause the one or more computers to:
 determine accuracy of the at least one model based on a confidence value associated with output of the at least one model; and   update the at least one training data set based on the confidence value associated with the output of the at least one model.   
     
     
         42 . A computer-implemented method, comprising:
 calculating, by a processor, a set of confidence values for a set of features associated with multi-dimensional data, wherein each confidence value of the set of confidence values is associated with an operator estimate, wherein each confidence value represents a probability that a feature representation of the multi-dimensional data belongs to a distribution, and wherein the operator estimate is associated with an operator determined by a statistical model;   updating, by the processor, at least one training data set with the multi-dimensional data in response to a determination, based on the set of confidence values, that the multi-dimensional data is included in the at least one training data set; and   generating, by the processor, at least one model from the at least one training data set.   
     
     
         43 . The computer-implemented method of  claim 42 , further comprising:
 calculating, by the processor, an operator estimation score for the multi-dimensional data based on the feature representation of the multi-dimensional data.   
     
     
         44 . The computer-implemented method of  claim 42 , further comprising:
 calculating, by the processor, an operator estimation score for the multi-dimensional data based on an estimator trained using the set of confidence values.   
     
     
         45 . The computer-implemented method of  claim 42 , further comprising:
 calculating, by the processor, a set of global estimate confidence values, wherein global estimate confidence value from the set of global estimate confidence values represents a probability of the feature representation belonging to a global distribution represented by a global data set that contains one or more data instances of a same type as the multi-dimensional data.   
     
     
         46 . The computer-implemented method of  claim 45 , further comprising:
 calculating, by the processor, a global estimation score based on the set of global estimate confidence values.   
     
     
         47 . The computer-implemented method of  claim 42 , further comprising:
 calculating, by the processor, an estimation score for the multi-dimensional data based on the set of confidence values; and   assigning, by the processor, the multi-dimensional data to a class associated with the at least one model based on the estimation score for the multi-dimensional data.   
     
     
         48 . The computer-implemented method of  claim 42 , further comprising:
 determining, by the processor, accuracy of the at least one model based on a confidence value associated with output of the at least one model; and   updating, by the processor, the at least one training data set based on the confidence value associated with the output of the at least one model.   
     
     
         49 . A computer program product, stored on a computer readable medium,
 comprising instructions that when executed by one or more computers cause the one or more computers to:   calculate a set of confidence values for a set of features associated with multi-dimensional data, wherein each confidence value from the set of confidence values is associated with an operator estimate, wherein each confidence value represents a probability that a feature representation of the multi-dimensional data belongs to a distribution, and wherein the operator estimate is associated with an operator determined by a statistical model;   update at least one training data set with the multi-dimensional data in response to a determination, based on the set of confidence values, that the multi-dimensional data is included in the at least one training data set; and   generate at least one model from the at least one training data set.   
     
     
         50 . The computer program product of  claim 49 , wherein the instructions, when executed by the one or more computers, further cause the one or more computers to:
 calculate an operator estimation score for the multi-dimensional data based on an estimator trained using the set of confidence values.   
     
     
         51 . The computer program product of  claim 49 , wherein the instructions, when executed by the one or more computers, further cause the one or more computers to:
 calculate a set of global estimate confidence values, wherein global estimate confidence value from the set of global estimate confidence values represents a probability of the feature representation belonging to a global distribution represented by a global data set that contains one or more data instances of a same type as the multi-dimensional data.   
     
     
         52 . The computer program product of  claim 51 , wherein the instructions, when executed by the one or more computers, further cause the one or more computers to:
 calculate a global estimation score based on the set of global estimate confidence values.

Join the waitlist — get patent alerts

Track US2020302337A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.