US2019378044A1PendingUtilityA1

Processing dynamic data within an adaptive oracle-trained learning system using curated training data for incremental re-training of a predictive model

Assignee: GROUPON INCPriority: Dec 23, 2013Filed: Aug 21, 2019Published: Dec 12, 2019
Est. expiryDec 23, 2033(~7.4 yrs left)· nominal 20-yr term from priority
G06N 20/00G06N 5/04
57
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In general, embodiments of the present invention provide systems, methods and computer readable media for an adaptive oracle-trained learning framework for automatically building and maintaining models that are developed using machine learning algorithms. In embodiments, the framework leverages at least one oracle (e.g., a crowd) for automatic generation of high-quality training data to use in deriving a model. Once a model is trained, the framework monitors the performance of the model and, in embodiments, leverages active learning and the oracle to generate feedback about the changing data for modifying training data sets while maintaining data quality to enable incremental adaptation of the model.

Claims

exact text as granted — not AI-modified
1 . A system comprising one or more computers configured to implement an adaptive learning framework for automatically building and maintaining a predictive model for processing dynamic data, wherein the adaptive learning framework is configured to include:
 the predictive model, wherein the model is configured to generate model output from processing an input data instance received by the adaptive learning framework, and wherein the model output includes a judgment and a confidence value representing certainty of the judgment;   a training data set from which the predictive model is derived using machine learning; and   a training data manager, wherein the training data manager is configured for curating the training data set; and   a labeled data reservoir configured to store labeled data instances that have been processed by the predictive model, wherein the labeled data reservoir includes a pool of possible training data, wherein the set of labeled data instances are not included in the training data; and   wherein each labeled data instance is associated with a true label representing the instance; and   wherein the training data manager is configured to perform operations comprising:   determining whether to update the training data set;   in an instance in which the training data set is to be updated,   selecting a set of labeled data instances from a labeled data reservoir; and   updating the training data using the set of labeled data instances.   
     
     
         2 . The system of  claim 1 , wherein determining whether to update the training data set is based at least in part on analyzing the distribution and quality of the training data. 
     
     
         3 . The system of  claim 1 , wherein determining whether to update the training data set is based at least in part on an accuracy assessment of the model performance. 
     
     
         4 . The system of  claim 3 , wherein the accuracy assessment is based on determining whether the confidence value satisfies a confidence threshold value. 
     
     
         5 . The system of  claim 2 , wherein the current model is a classifier predicting to which of a set of predictive categories an input data instance belongs, wherein a true label associated with a labeled data instance identifies the predictive category to which the labeled data instance belongs, and wherein selecting the set of labeled data instances from the labeled data reservoir is based at least in part on maintaining a class balance within the training data. 
     
     
         6 . The system of  claim 1 , wherein the labeled data reservoir includes labeled data instances that are received from multiple sources, and wherein selecting a labeled data instance from the set of labeled data instances comprises:
 comparing a source of the labeled data instance with a pre-determined source; and   selecting the labeled data instance in an instance in which the source of the labeled data instance matches the pre-determined source.   
     
     
         7 . The system of  claim 1 , wherein the operations further comprise:
 in response to updating the training data, determining whether to re-train the model;   in an instance in which the model is re-trained,   generating at least one candidate training data set using the updated training data;   deriving a candidate model using the candidate training data set;   generating an assessment of whether the candidate model performance is improved from the model performance; and   instantiating the candidate training data set and the candidate model in the adaptive learning framework in an instance in which the candidate model performance is improved from the model performance.   
     
     
         8 . The system of  claim 7 , wherein generating the assessment of whether the candidate model performance is improved from the model performance includes A/B testing. 
     
     
         9 . The system of  claim 8 , wherein generating the assessment comprises calculating a cross-validation between the candidate model performance and the model performance. 
     
     
         10 . The system of  claim 8 , wherein there are multiple candidate models, and wherein generating the assessment respectively for each of the multiple candidate models is implemented in parallel. 
     
     
         11 . A computer program product, stored on a non-transitory computer readable medium, comprising instructions that when executed on one or more computers cause the one or more computers to implement an adaptive learning framework for automatically building and maintaining a predictive model for processing dynamic data, wherein the adaptive learning framework is configured to include:
 the predictive model, wherein the model is configured to generate model output from processing an input data instance received by the adaptive learning framework, and wherein the model output includes a judgment and a confidence value representing certainty of the judgment;   a training data set from which the predictive model is derived using machine learning; and   a training data manager, wherein the training data manager is configured for curating the training data set; and   a labeled data reservoir configured to store labeled data instances that have been processed by the predictive model, wherein the labeled data reservoir includes a pool of possible training data, wherein the set of labeled data instances are not included in the training data; and   wherein each labeled data instance is associated with a true label representing the instance; and   wherein the training data manager is configured to perform operations comprising:   determining whether to update the training data set;   in an instance in which the training data set is to be updated,   selecting a set of labeled data instances from a labeled data reservoir; and   updating the training data using the set of labeled data instances.   
     
     
         12 . The computer program product of  claim 11 , wherein determining whether to update the training data set is based at least in part on analyzing the distribution and quality of the training data. 
     
     
         13 . The computer program product of  claim 11 , wherein determining whether to update the training data set is based at least in part on an accuracy assessment of the model performance. 
     
     
         14 . The computer program product of  claim 13 , wherein the accuracy assessment is based on determining whether the confidence value satisfies a confidence threshold value. 
     
     
         15 . The computer program product of  claim 12 , wherein the current model is a classifier predicting to which of a set of predictive categories an input data instance belongs, wherein a true label associated with a labeled data instance identifies the predictive category to which the labeled data instance belongs, and wherein selecting the set of labeled data instances from the labeled data reservoir is based at least in part on maintaining a class balance within the training data. 
     
     
         16 . The computer program product of  claim 11 , wherein the labeled data reservoir includes labeled data instances that are received from multiple sources, and wherein selecting a labeled data instance from the set of labeled data instances comprises:
 comparing a source of the labeled data instance with a pre-determined source; and   selecting the labeled data instance in an instance in which the source of the labeled data instance matches the pre-determined source.   
     
     
         17 . The computer program product of  claim 11 , wherein the operations further comprise:
 in response to updating the training data, determining whether to re-train the model;   in an instance in which the model is re-trained,   generating at least one candidate training data set using the updated training data;   deriving a candidate model using the candidate training data set;   generating an assessment of whether the candidate model performance is improved from the model performance; and   instantiating the candidate training data set and the candidate model in the adaptive learning framework in an instance in which the candidate model performance is improved from the model performance.   
     
     
         18 . The computer program product of  claim 17 , wherein generating the assessment of whether the candidate model performance is improved from the model performance includes A/B testing. 
     
     
         19 . The computer program product of  claim 18 , wherein generating the assessment comprises calculating a cross-validation between the candidate model performance and the model performance. 
     
     
         20 . The computer program product of  claim 18 , wherein there are multiple candidate models, and wherein generating the assessment respectively for each of the multiple candidate models is implemented in parallel.

Join the waitlist — get patent alerts

Track US2019378044A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.