US2022156646A1PendingUtilityA1

Rejecting Biased Data Using A Machine Learning Model

Assignee: GOOGLE LLCPriority: Sep 10, 2018Filed: Jan 31, 2022Published: May 19, 2022
Est. expirySep 10, 2038(~12.1 yrs left)· nominal 20-yr term from priority
G06F 18/27G06F 18/24G06N 5/04G06F 18/2321G06F 18/22G06N 20/00
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method for rejecting biased data using a machine learning model includes receiving a cluster training data set including a known unbiased population of data and training a clustering model to segment the received cluster training data set into clusters based on data characteristics of the known unbiased population of data. Each cluster of the cluster training data set includes a cluster weight. The method also includes receiving a training data set for a machine learning model and generating training data set weights corresponding to the training data set for the machine learning model based on the clustering model. The method also includes adjusting each training data set weight of the training data set weights to match a respective cluster weight and providing the adjusted training data set to the machine learning model as an unbiased training data set.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:
 receiving a training data set for training a machine learning model, the training data set comprising one or more bias sensitive variables;   generating, using a bias scoring model, a bias score for the training data set;   determining whether the bias score satisfies a threshold score;   when the bias score satisfies the threshold score, labeling the training set with an approval indicator; and   when the bias score fails to satisfy the threshold score, labeling the training set with a rejection indicator.   
     
     
         2 . The method of  claim 1 , wherein the operations further comprise training the bias scoring model using one or more bias scoring training data sets. 
     
     
         3 . The method of  claim 2 , wherein the bias scoring model is trained using feedback based on the bias score. 
     
     
         4 . The method of  claim 1 , wherein the bias score is a numerical representation of bias. 
     
     
         5 . The method of  claim 1 , wherein the threshold score comprises a user-configurable acceptable bias score. 
     
     
         6 . The method of  claim 1 , wherein the operations further comprise, when the bias score satisfies the threshold score, training the machine learning model using the training data set. 
     
     
         7 . The method of  claim 1 , wherein the operations further comprise, when the bias score fails to satisfy the threshold score, generating, using a bias rejection model, an unbiased training data set from the training data set. 
     
     
         8 . The method of  claim 7 , wherein the operations further comprise, when the bias score fails to satisfy the threshold score, training the machine learning model using the unbiased training data set. 
     
     
         9 . The method of  claim 7 , wherein generating the unbiased training data set comprises:
 segmenting the training data set into a plurality of clusters, wherein at least one cluster of the plurality of clusters corresponds to a bias sensitive variable of the one or more bias sensitive variables;   generating a weight for each cluster of the plurality of clusters; and   generating the unbiased training data set based on the weighted plurality of clusters.   
     
     
         10 . The method of  claim 9 , wherein the weight for each cluster is based on a probability distribution of data characteristics of a target population. 
     
     
         11 . A system comprising:
 data processing hardware; and   memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
 receiving a training data set for training a machine learning model, the training data set comprising one or more bias sensitive variables; 
 generating, using a bias scoring model, a bias score for the training data set; 
 determining whether the bias score satisfies a threshold score; 
 when the bias score satisfies the threshold score, labeling the training set with an approval indicator; and 
 when the bias score fails to satisfy the threshold score, labeling the training set with a rejection indicator. 
   
     
     
         12 . The system of  claim 11 , wherein the operations further comprise training the bias scoring model using one or more bias scoring training data sets. 
     
     
         13 . The system of  claim 12 , wherein the bias scoring model is trained using feedback based on the bias score. 
     
     
         14 . The system of  claim 11 , wherein the bias score is a numerical representation of bias. 
     
     
         15 . The system of  claim 11 , wherein the threshold score comprises a user-configurable acceptable bias score. 
     
     
         16 . The system of  claim 11 , wherein the operations further comprise, when the bias score satisfies the threshold score, training the machine learning model using the training data set. 
     
     
         17 . The system of  claim 11 , wherein the operations further comprise, when the bias score fails to satisfy the threshold score, generating, using a bias rejection model, an unbiased training data set from the training data set. 
     
     
         18 . The system of  claim 17 , wherein the operations further comprise, when the bias score fails to satisfy the threshold score, training the machine learning model using the unbiased training data set. 
     
     
         19 . The system of  claim 17 , wherein generating the unbiased training data set comprises:
 segmenting the training data set into a plurality of clusters, wherein at least one cluster of the plurality of clusters corresponds to a bias sensitive variable of the one or more bias sensitive variables;   generating a weight for each cluster of the plurality of clusters; and   generating the unbiased training data set based on the weighted plurality of clusters.   
     
     
         20 . The system of  claim 19 , wherein the weight for each cluster is based on a probability distribution of data characteristics of a target population.

Join the waitlist — get patent alerts

Track US2022156646A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.