US2021073683A1PendingUtilityA1

Machine learning models for evaluating differences between groups and methods thereof

Assignee: CAPITAL ONE SERVICES LLCPriority: Dec 12, 2018Filed: Nov 16, 2020Published: Mar 11, 2021
Est. expiryDec 12, 2038(~12.4 yrs left)· nominal 20-yr term from priority
G06T 11/26G06N 5/01G06N 20/20G06N 5/04G06N 20/00G06T 11/206
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems, methods, and computer readable media are disclosed for generating, modifying, and using machine learning models to predict and evaluate differences between groups. Methods disclosed herein may include identifying variables that characterize members of a first group, generating shift indicators using the identified variables, generating a machine learning model using the shift indicators and the first group, using the machine learning model and the group to predict shifts between the first group and a predicted second group, determining an aggregate population shift and an aggregate performance shift between the first group and an actual second group, and identifying an impact of one or more of the shift indicators on the aggregate population shift or performance shift. Systems and methods disclosed herein may be configured to receive requests to predict and evaluate differences between group, and to return such predictions and evaluations to one or more users.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled) 
     
     
         21 . A computer-implemented method for generating a machine learning model to define differences between groups, the method comprising:
 retrieving a plurality of first variables associated with members of a first group, wherein the plurality of first variables define characteristics of the members of the first group;   transforming the plurality of first variables into one or more shift indicators based on characteristic similarities between the members of the first group, wherein a quantity of the one or more shift indicators is less than the plurality of first variables;   training the machine learning model to predict members of a target group by:
 generating members of a predicted second group using the one or more shift indicators; 
 determining a predicted population shift between the members of the first group and the members of the predicted second group; 
 retrieving a plurality of second variables associated with members of an actual second group, wherein the plurality of second variables define characteristics of the members of the actual second group; 
 determining an actual population shift representing a change in population between the first group and the actual second group based on a comparison of the plurality of first variables and the plurality of second variables; 
 determining an aggregate population shift based on a difference between the predicted population shift and the actual population shift; and 
 determining at least one of the plurality of first variables has an impact on the aggregate population shift that is greater than a remaining plurality of first variables. 
   
     
     
         22 . The computer-implemented method of  claim 21 , further comprising:
 outputting an identification of the at least one of the plurality of first variables having the impact on the aggregate population shift that is greater than the remaining plurality of first variables.   
     
     
         23 . The computer-implemented method of  claim 21 , wherein prior to determining the at least one of the plurality of first variables has the impact on the aggregate population shift that is greater than the remaining plurality of first variables, the method comprises:
 determining at least one shift indicator has the impact on the aggregate population shift that is greater than a remaining one or more shift indicators; and   determining the at least one shift indicator is associated with a subset of the plurality of first variables that includes the at least one of the plurality of first variables.   
     
     
         24 . The computer-implemented method of  claim 23 , wherein determining the at least one shift indicator has the impact on the aggregate population shift that is greater than the remaining one or more shift indicators comprises:
 determining a coefficient of impact of the one or more shift indicators transformed from the plurality of first variables;   comparing the coefficient of impact of each of the one or more shift indicators to one another; and   determining the coefficient of impact of the at least one shift indicator is greater than the coefficient of impact of the remaining one or more shift indicators.   
     
     
         25 . The computer-implemented method of  claim 24 , further comprising:
 outputting the coefficient of impact of the at least one shift indicator with the impact on the aggregate population shift that is greater than the remaining one or more shift indicators.   
     
     
         26 . The computer-implemented method of  claim 23 , further comprising:
 determining the impact of the subset of the plurality of first variables associated with the at least one shift indicator; and   determining the at least one of the plurality of first variables has the impact that is greater than a remainder of the subset of the plurality of first variables.   
     
     
         27 . The computer-implemented method of  claim 21 , wherein prior to determining the at least one of the plurality of first variables has the impact on the aggregate population shift that is greater than the remaining plurality of first variables, the method comprises:
 determining an impact level of each of the plurality of first variables on the aggregate population shift.   
     
     
         28 . The computer-implemented method of  claim 27 , further comprising:
 ranking the plurality of first variables relative to one another based on the impact level on the aggregate population shift.   
     
     
         29 . The computer-implemented method of  claim 28 , further comprising:
 generating a waterfall chart that includes representations of the ranking of the plurality of first variables based on the impact level on the aggregate population shift; and   outputting the waterfall chart on a display.   
     
     
         30 . The computer-implemented method of  claim 21 , wherein the members of the first group comprise a first population exhibiting a first loss of monetary assets over a first period of time; and
 wherein the plurality of first variables associated with the members of the first group comprises loss percentages associated with loss of monetary assets.   
     
     
         31 . The computer-implemented method of  claim 30 , wherein the members of the predicted second group comprise a predicted second population exhibiting a predicted second loss of monetary assets over a second period of time; and
 wherein the plurality of second variables associated with the members of the actual second group comprises loss percentages associated with loss of monetary assets.   
     
     
         32 . The computer-implemented method of  claim 21 , wherein transforming the plurality of first variables into the one or more shift indicators further comprises:
 transforming a variable value of each of the plurality of first variables into a binary value corresponding to the one or more shift indicators.   
     
     
         33 . The computer-implemented method of  claim 21 , wherein the machine learning model is one of a gradient boosting model or a random forest model. 
     
     
         34 . A computer-implemented method for generating a machine learning model to define differences between groups, the method comprising:
 identifying a plurality of variables that characterize each member of a first group and a second group;   transforming the plurality of variables that characterize members of the first group into a plurality of shift indicators, wherein each of the plurality of shift indicators include a subset of the plurality of variables, such that the plurality of variables exceeds the plurality of shift indicators;   training the machine learning model, using the plurality of shift indicators, to predict members of a target group that are different than the first group and the second group by:
 predicting a predicted second group based on the plurality of shift indicators and the first group; 
 predicting a predicted population shift between the predicted second group and the first group; 
 calculating an actual population shift between the first group and the second group based on a difference between the plurality of variables that characterize each member of the first group and each member of the second group; 
 calculating an aggregate population shift between the actual population shift and the predicted population shift; and 
 determining a first variable of the plurality of variables having a greatest impact on the aggregate population shift. 
   
     
     
         35 . The computer-implemented method of  claim 34 , further comprising:
 determining a coefficient of impact on the aggregate population shift for each of the plurality of shift indicators;   multiplying the aggregate population shift by the coefficient of impact on the aggregate population shift for each of the plurality of shift indicators; and   determining an impact level of each of the plurality of shift indicators on the aggregate population shift based on the multiplication of the aggregate population shift and the coefficient of impact.   
     
     
         36 . The computer-implemented method of  claim 34 , wherein prior to determining the first variable of the plurality of variables has the greatest impact on the aggregate population shift, the method comprises:
 determining a first shift indicator of the plurality of shift indicators has the greatest impact on the aggregate population shift; and   determining the first shift indicator is associated with the first variable.   
     
     
         37 . The computer-implemented method of  claim 34 , wherein determining the first variable of the plurality of variables has the greatest impact on the aggregate population shift further comprises:
 determining at least one of the plurality of shift indicators has the greatest impact on the aggregate population shift; and   identifying the subset of the plurality of variables associated with the at least one of the plurality of shift indicators having the greatest impact on the aggregate population shift; and   determining the one or more of the plurality of variables in the subset include the first variable having the greatest impact on the aggregate population shift.   
     
     
         38 . The computer-implemented method of  claim 34 , further comprising:
 determining an impact level of each of the plurality of variables on the aggregate population shift; and   ranking the plurality of variables relative to one another based on the impact level on the aggregate population shift.   
     
     
         39 . The computer-implemented method of  claim 38 , further comprising:
 generating a waterfall chart that includes representations of the ranking of the plurality of variables relative to one another based on the impact level on the aggregate population shift; and   outputting the waterfall chart on a display.   
     
     
         40 . A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a computer system, cause the one or more processors to perform operations comprising:
 identifying a plurality of first variable values that characterize members of a first group;   identifying a plurality of second variable values that characterize members of a second group;   determining a plurality of shift indicators based on the plurality of first variable values, wherein each of the plurality of shift indicators includes a subset of the plurality of first variable values having similar values to one another, wherein the plurality of shift indicators is fewer in number than the plurality of first variable values;   generating a machine learning model, using the plurality of shift indicators and the first group, to predict members of a target group by:
 predicting members of a predicted second group; 
 predicting a predicted population shift between the members of the first group and the members of the predicted second group; 
 determining a first difference between the members of the first group and the members of the second group, based on the plurality of first variable values and the plurality of second variable values, to calculate an actual population shift; 
 determining a second difference between the predicted population shift and the actual population shift to calculate an aggregate population shift; 
 determining an impact of each of the plurality of first variable values on the aggregate population shift; and 
 generating a ranking of the plurality of first variable values relative to one another based on the impact of each on the aggregate population shift.

Join the waitlist — get patent alerts

Track US2021073683A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.