Machine learning model development and optimization process that ensures performance validation and data sufficiency for regulatory approval
Abstract
Machine learning model development and optimization tools are provided that ensure performance validation and data sufficiency for regulatory approval. According to an embodiment, a computer implemented method can comprise training a machine learning model to perform an inferencing task on an initial set of data samples included in a sample population. In various embodiments, the model can include a medical AI model. The method further comprises determining, by the system, subgroup performance measures for subgroups of the data samples respectively associated with different metadata factors, wherein the subgroup performance measures reflect performance accuracy of the machine learning model with respect to the subgroups. The method further comprises determining, by the system, whether the machine learning model meets an acceptable level of performance for deployment in a field environment based on whether the subgroup performance measures respectively satisfy a threshold subgroup performance measure.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system, comprising:
a memory that stores computer executable components; and a processor that executes the computer executable components stored in the memory, wherein the computer executable components comprise:
a performance evaluation component evaluates performance of a machine learning model trained to perform an inferencing task regarding an assessment of medical data samples, wherein the performance evaluation component determines subgroup performance measures for different subgroups of the medical data samples grouped based on different metadata factors comprising clinical and non-clinical metadata factors, and wherein the subgroup performance measures reflect measures of performance accuracy of the machine learning model with respect to the different subgroups; and
an approval regulation component that determines whether the machine learning model meets an acceptable level of performance for deployment in a field environment based on whether the subgroup performance measures respectively satisfy a threshold subgroup performance measure.
2 . The system of claim 1 , wherein the subgroup performance measures respectively comprise uncertainty estimate values representative of a degree of uncertainty in the performance accuracy of the machine learning model with respect to the different subgroups, and wherein the threshold subgroup performance measure comprises a maximum uncertainty value.
3 . The system of claim 1 , wherein the subgroup performance measures respectively comprise lower prediction bound values of the performance accuracy of the machine learning model with respect to the different subgroups, and wherein the threshold subgroup performance measure comprises a minimum lower prediction bound value.
4 . The system of claim 1 , wherein based on a determination that a subgroup of the different subgroups has a subgroup performance measure that fails to satisfy the threshold subgroup performance measure, the approval regulation component disapproves the machine learning model as having the acceptable level of performance for deployment in the field environment on new data samples included in the subgroup.
5 . The system of claim 1 , wherein based on a determination that all of the subgroup the subgroup performance measures satisfy the threshold subgroup performance measure, the approval regulation component approves the machine learning model as having the acceptable level of performance for deployment in the field environment.
6 . The system of claim 1 , wherein the computer executable components further comprise:
an active learning component that identifies underperforming subgroups of the different subgroups of the medical data samples based on respective subgroup performance measures associated with the underperforming subgroups failing to satisfy the threshold subgroup performance measure; and an active sampling component that retrieves additional data samples for the underperforming subgroups from a collection of population data samples and adds the additional data samples to at least one of a training dataset, a test dataset, a validation dataset or a regulatory validation dataset.
7 . The system of claim 6 , wherein the computer executable components further comprise:
a training component the updates the machine learning model using at least one of the training dataset, the test dataset, the validation dataset or the regulatory validation dataset, resulting in an updated machine learning model, and wherein the performance evaluation component further updates the respective subgroup performance measures based on new measures of performance accuracy of the updated machine learning model with respect to the underperforming subgroups.
8 . The system of claim 7 , wherein the active sampling component continues to retrieve the additional data samples and the model training component continues to train, update and validate the machine learning model using the additional data samples until all of the subgroup performance measures respectively satisfy the threshold subgroup performance measure or a maximum amount, by cost or count, of the additional data samples authorized for retrieval has been reached.
9 . The system of claim 6 , wherein the active sampling component further determines a difficulty score for the underperforming subgroups, and wherein the active sampling component further selects the additional data samples that maximize a change to the difficulty score.
10 . The system of claim 6 , wherein the active sampling component determines priority scores for potential new data samples based on the respective subgroup performance measures of the underperforming subgroups that the potential new data samples respectively belong, and wherein the active sampling component further selects the additional data samples from the potential new data samples based on the priority scores.
11 . The system of claim 6 , wherein the active sampling component further determines an amount of the additional data samples to retrieve using an entitlement function, including first amount of additional training data samples of the additional data samples and a second amount of additional validation data samples of the additional data samples.
12 . A method, comprising:
evaluating, by a system operatively coupled to a processor, performance of a machine learning model trained to perform an inferencing task regarding an assessment of medical data samples, wherein the evaluating comprises determining subgroup performance measures for different subgroups of the medical data samples grouped based on different metadata factors comprising clinical and non-clinical metadata factors, and wherein the subgroup performance measures reflect measures of performance accuracy of the machine learning model with respect to the different subgroups; and determining, by the system, whether the machine learning model meets an acceptable level of performance for deployment in a field environment based on whether the subgroup performance measures respectively satisfy a threshold subgroup performance measure.
13 . The method of claim 12 , wherein the subgroup performance measures respectively comprise uncertainty estimate values representative of a degree of uncertainty in the performance accuracy of the machine learning model with respect to the different subgroups, and wherein the threshold subgroup performance measure comprises a maximum uncertainty value.
14 . The method of claim 12 , wherein based on a determination that a subgroup of the different subgroups has a subgroup performance measure that fails to satisfy the threshold subgroup performance measure, the method further comprises:
disapproving, by the system, the machine learning model as having the acceptable level of performance for deployment in the field environment on new data samples included in the subgroup.
15 . The method of claim 12 , wherein based on a determination that all of the subgroup the subgroup performance measures satisfy the threshold subgroup performance measure, the method further comprises:
approving, by the system, the machine learning model as having the acceptable level of performance for deployment in the field environment.
16 . The method of claim 12 , further comprising:
identifying, by the system, underperforming subgroups of the different subgroups of the medical data samples based on respective subgroup performance measures associated with the underperforming subgroups failing to satisfy the threshold subgroup performance measure; and retrieving, by the system, additional data samples for the underperforming subgroups from a collection of population data samples and adds the additional data samples to at least one of a training dataset, a test dataset, a validation dataset or a regulatory validation dataset.
17 . The method of claim 16 , further comprising:
training, by the system, the machine learning model using at least one of the training dataset, the test dataset, the validation dataset or the regulatory validation dataset, resulting in an updated machine learning model; and updating, by the system, the respective subgroup performance measures based on new measures of performance accuracy of the updated machine learning model with respect to the underperforming subgroups.
18 . The method of claim 16 , further comprising:
continuing, by the system, the retrieving, the training and the updating until all of the subgroup performance measures respectively satisfy the threshold subgroup performance measure or a maximum amount, by cost or count, of the additional data samples authorized for retrieval has been reached.
19 . The method of claim 16 , further comprising:
determining, by the system, a difficulty score for the underperforming subgroups; and selecting, by the system, the additional data samples that maximize a change to the difficulty score.
20 . A machine-readable storage medium, comprising executable instructions that, when executed by a processor, facilitate performance of operations, comprising:
evaluating performance of a machine learning model trained to perform an inferencing task regarding an assessment of medical data samples, wherein the evaluating comprises determining subgroup performance measures for different subgroups of the medical data samples grouped based on different metadata factors comprising clinical and non-clinical metadata factors, and wherein the subgroup performance measures reflect measures of performance accuracy of the machine learning model with respect to the different subgroups; and determining whether the machine learning model meets an acceptable level of performance for deployment in a field environment based on whether the subgroup performance measures respectively satisfy a threshold subgroup performance measure.Join the waitlist — get patent alerts
Track US2023229972A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.