US2025371437A1PendingUtilityA1

Training predictive models based on reward signals and hyperparameter searching

Assignee: INTUIT INCPriority: May 31, 2024Filed: May 31, 2024Published: Dec 4, 2025
Est. expiryMay 31, 2044(~17.8 yrs left)· nominal 20-yr term from priority
G06N 5/01G06N 20/00G06N 20/20
59
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Aspects of the present disclosure provide techniques for training and using machine learning models to predict and present an optimal workflow to a user of a software application. An example method generally includes generating a training data set including a plurality of exemplars including features associated a user of a software application, a sequence of workflow steps presented to the user of the software application, and a reward metric. A plurality of hyperparameter sets for training a plurality of predictive models is generated. The plurality of predictive models are trained based on the plurality of hyperparameter sets. A hyperparameter set from the plurality of hyperparameter sets is selected based on performance metrics for each of the plurality of predictive models. A machine learning model is trained based on the selected hyperparameter set and the training data set, and the trained machine learning model is deployed.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A processor-implemented method, comprising:
 generating a training data set including a plurality of exemplars, each exemplar including features associated a user of a software application, a sequence of workflow steps presented to the user of the software application, and a reward metric associated with the user of the software application and the sequence of workflow steps;   generating a plurality of hyperparameter sets for training a plurality of predictive models for identifying a sequence of workflow steps to present to users of the software application;   training the plurality of predictive models based on the plurality of hyperparameter sets;   selecting a hyperparameter set from the plurality of hyperparameter sets based on performance metrics for each of the plurality of predictive models;   training a machine learning model based on the selected hyperparameter set and the training data set; and   deploying the machine learning model.   
     
     
         2 . The method of  claim 1 , wherein generating the plurality of hyperparameter sets comprises generating, for each respective hyperparameter set of the plurality of hyperparameter sets, a respective random seed value for separating the training data set into a training subset and a validation subset. 
     
     
         3 . The method of  claim 1 , wherein:
 the predictive models comprise tree-based uplift models, and   each hyperparameter set of the plurality of hyperparameter sets comprise a tree depth parameter, a number of trees parameter, and a tree split value parameter.   
     
     
         4 . The method of  claim 1 , wherein training the plurality of predictive models comprises training a model to maximize the reward metric. 
     
     
         5 . The method of  claim 1 , wherein selecting the hyperparameter set from the plurality of hyperparameter sets comprises calculating, for each respective predictive model of the plurality of predictive models, one or more error metrics between an average reward over a plurality of sequences of workflow steps and an average reward for outputs of the respective predictive model matching a defined baseline policy. 
     
     
         6 . The method of  claim 5 , wherein selecting the hyperparameter set further comprises discarding hyperparameter sets associated with predictive models having a predictive value below a threshold value with at least one error metric of the one or more error metrics being a negative value. 
     
     
         7 . The method of  claim 1 , wherein selecting the hyperparameter set comprises selecting the hyperparameter set having a ratio of performant models to total models exceeding a threshold value and no nonperformant models. 
     
     
         8 . A processor-implemented method, comprising:
 receiving, from a user of a software application, a request to execute a workflow in the software application;   generating, using a predictive model and features associated with the user of the software application, a workflow sequence that maximizes a reward metric for the user of the software application, the predictive model comprising one of a plurality of predictive models trained based on a set of hyperparameters resulting in a set of predictive models having a highest level of performance; and   executing the generated workflow sequence.   
     
     
         9 . The method of  claim 8 , wherein the features associated with the user of the software application comprise at least one of static features defining characteristics of the user of the software application or dynamic features associated with user activity within the software application. 
     
     
         10 . The method of  claim 8 , wherein the reward metric comprises a cumulative reward metric calculated over each step in the generated workflow sequence. 
     
     
         11 . The method of  claim 8 , wherein the reward metric comprises a total revenue associated with completion of the workflow. 
     
     
         12 . A processing system, comprising:
 at least one memory having executable instructions stored thereon; and   one or more processors configured to execute the executable instructions to cause the processing system to:
 generate a training data set including a plurality of exemplars, each exemplar including features associated a user of a software application, a sequence of workflow steps presented to the user of the software application, and a reward metric associated with the user of the software application and the sequence of workflow steps; 
 generate a plurality of hyperparameter sets for training a plurality of predictive models for identifying a sequence of workflow steps to present to users of the software application; 
 train the plurality of predictive models based on the plurality of hyperparameter sets; 
 select a hyperparameter set from the plurality of hyperparameter sets based on performance metrics for each of the plurality of predictive models; 
 train a machine learning model based on the selected hyperparameter set and the training data set; and 
 deploy the machine learning model. 
   
     
     
         13 . The processing system of  claim 12 , wherein to generate the plurality of hyperparameter sets, the one or more processors are configured to cause the processing system to generate, for each respective hyperparameter set of the plurality of hyperparameter sets, a respective random seed value for separating the training data set into a training subset and a validation subset. 
     
     
         14 . The processing system of  claim 12 , wherein:
 the predictive models comprise tree-based uplift models, and   each hyperparameter set of the plurality of hyperparameter sets comprise a tree depth parameter, a number of trees parameter, and a tree split value parameter.   
     
     
         15 . rocessing system of  claim 12 , wherein to train the plurality of predictive models, the one or more processors are configured to cause the processing system to train a model to maximize the reward metric. 
     
     
         16 . The processing system of  claim 12 , wherein to select the hyperparameter set from the plurality of hyperparameter sets, the one or more processors are configured to cause the processing system to calculate, for each respective predictive model of the plurality of predictive models, one or more error metrics between an average reward over a plurality of sequences of workflow steps and an average reward for outputs of the respective predictive model matching a defined baseline policy. 
     
     
         17 . The processing system of  claim 16 , wherein to select the hyperparameter set, the one or more processors are further configured to cause the processing system to discard hyperparameter sets associated with predictive models having a predictive value below a threshold value with at least one error metric of the one or more error metrics being a negative value. 
     
     
         18 . The processing system of  claim 12 , wherein to select the hyperparameter set, the one or more processors are configured to cause the processing system to select the hyperparameter set having a ratio of performant models to total models exceeding a threshold value and no nonperformant models.

Join the waitlist — get patent alerts

Track US2025371437A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.