System and method for generating competing models in rashomon sets for gradient boosting
Abstract
A method may include: receiving a dataset comprising a plurality of samples and a loss function; training a first number first machine learning models using the dataset comprising, wherein each of the first machine learning models has a similar performance; selecting one of the first machine learning models with a smallest loss; computing a residual for each of the plurality of samples using the one first machine learning model; defining a new dataset comprising the plurality of samples and the residual for each samples; training the first machine learning model with the new dataset; generating a second plurality of machine learning models by repeating the selecting, the computing, the defining, and training for a number of boosting iterations; selecting a subset of the second plurality of machine learning model models having a specified property; and deploying the subset of second machine learning models to a downstream task.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
receiving, by a computer program, a dataset comprising a plurality of samples and a loss function; training, by the computer program, a first number of a plurality of first machine learning models using the dataset, wherein each of the plurality of first machine learning models has a similar performance as measured by the loss function; selecting, by the computer program, one of the first machine learning models with a smallest loss using the loss function; computing, by the computer program, a residual for each of the plurality of samples using the one first machine learning model; defining, by the computer program, a new dataset comprising the plurality of samples and the residual for each samples; training, by the computer program, the first machine learning model with the new dataset; generating, by the computer program, a second plurality of machine learning models by repeating the selecting, the computing, the defining, and training for a number of boosting iterations, wherein a number of second machine learning models is equal to the first number multiplied by the number of boosting iterations; selecting, by the computer program, a subset of the second plurality of machine learning model models having a specified property; and deploying, by the computer program, the subset of second machine learning models to a downstream task.
2 . The method of claim 1 , wherein the computer program further receives the number of boosting iterations.
3 . The method of claim 1 , further comprising:
receiving, by the computer program, a hypothesis space comprising one of sparse decision-trees, linear models, and neural networks.
4 . The method of claim 1 , wherein the first plurality of machine learning models are trained with different initializations or different random seeds to fit the dataset.
5 . The method of claim 1 , wherein the specified property comprises fairness, and fairness is measured using a statistical parity for the plurality of second machine learning models.
6 . The method of claim 1 , wherein the specified property comprises interpretability, and interpretability is measured using a SHapley Additive explanations value for each of the plurality of second machine learning models.
7 . The method of claim 1 , further complying:
computing, by the computer program, a predictive multiplicity metric for the second plurality of machine learning models, wherein the predictive multiplicity metric measures conflicting predictions among the second plurality of machine learning models.
8 . A non-transitory computer readable storage medium, including instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to perform steps comprising:
receiving a dataset comprising a plurality of samples and a loss function; training a first number of a plurality of first machine learning models using the dataset, wherein each of the plurality of first machine learning models has a similar performance as measured by the loss function; selecting one of the first machine learning model with a smallest loss using the loss function; computing a residual for each of the plurality of samples using the one first machine learning models; defining a new dataset comprising the plurality of samples and the residual for each samples; training the first machine learning model with the new dataset; generating a second plurality of machine learning models by repeating the selecting, the computing, the defining, and the training for a number of boosting iterations, wherein a number of second machine learning models is equal to the first number multiplied by the number of boosting iterations; selecting a subset of the second plurality of machine learning model models having a specified property; and deploying the subset of second machine learning models to a downstream task.
9 . The non-transitory computer readable storage medium of claim 8 , further including instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to receive the number of boosting iterations.
10 . The non-transitory computer readable storage medium of claim 8 , further including instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to perform steps comprising:
receiving a hypothesis space comprising one of sparse decision-trees, linear models, and neural networks.
11 . The non-transitory computer readable storage medium of claim 8 , wherein the first plurality of machine learning models are trained with different initializations or different random seeds to fit the dataset.
12 . The non-transitory computer readable storage medium of claim 8 , wherein the specified property comprises fairness, and fairness is measured using a statistical parity for the plurality of second machine learning models.
13 . The non-transitory computer readable storage medium of claim 8 , wherein the specified property comprises interpretability, and interpretability is measured using a SHapley Additive explanations value for each of the plurality of second machine learning models.
14 . The non-transitory computer readable storage medium of claim 8 , further including instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to perform steps comprising:
computing a predictive multiplicity metric for the second plurality of machine learning models, wherein the predictive multiplicity metric measures conflicting predictions among the second plurality of machine learning models.
15 . A system, comprising:
a database storing a dataset; a user electronic device; and an electronic device executing a computer program that is configured to receive the dataset from the database and a loss function from the user electronic device; to train a first number of a plurality of first machine learning models using a dataset comprising a plurality of samples with different initializations or different random seeds to fit the dataset, wherein each of the plurality of first machine learning models has a similar performance as measured by the loss function; to select one of the first machine learning model with a smallest loss using the loss function; to compute a residual for each of the plurality of samples using the one first machine learning model; to define a new dataset comprising the plurality of samples and the residual for each samples; to train the first machine learning model with the new dataset; to generate a second plurality of machine learning models by repeating the selecting, the computing, the defining, and the training for a number of boosting iterations, wherein a number of second machine learning models is equal to the first number multiplied by the number of boosting iterations; to select a subset of the second plurality of machine learning model models having a specified property; and to deploy the subset of second machine learning models to a downstream task.
16 . The system of claim 15 , wherein the computer program further receives the number of boosting iterations.
17 . The system of claim 15 , wherein the computer program is further configured to receive a hypothesis space comprising one of sparse decision-trees, linear models, and neural networks.
18 . The system of claim 15 , wherein the specified property comprises fairness, and fairness is measured using a statistical parity for the plurality of second machine learning models.
19 . The system of claim 15 , wherein the specified property comprises interpretability, and interpretability is measured using a SHapley Additive explanations value for each of the plurality of second machine learning models.
20 . The system of claim 15 , wherein the computer program is further configured to compute a predictive multiplicity metric for the second plurality of machine learning models, wherein the predictive multiplicity metric measures conflicting predictions among the second plurality of machine learning models.Join the waitlist — get patent alerts
Track US2025348794A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.