US2025307714A1PendingUtilityA1

Machine learning program, method, and apparatus

Assignee: FUJITSU LTDPriority: Feb 9, 2023Filed: Jun 16, 2025Published: Oct 2, 2025
Est. expiryFeb 9, 2043(~16.5 yrs left)· nominal 20-yr term from priority
Inventors:Ryosuke Sonoda
G06N 20/00
68
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A machine learning apparatus a processor that executes a procedure. The procedure includes: calculating an independence between a prediction result in a case in which each of a plurality of items of unlabeled data is input to a machine learning model, and a value of a first attribute of each of the plurality of items of data; selecting first data from the plurality of items of data based on the independence; acquiring a label of the first data; and executing training of the machine learning model based on the first data and the label.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A non-transitory recording medium storing a program executable by a computer to perform machine learning program processing, the processing comprising:
 calculating an independence between a prediction result in a case in which each of a plurality of items of unlabeled data is input to a machine learning model, and a value of a first attribute of each of the plurality of items of data;   selecting first data from the plurality of items of data based on the independence;   acquiring a label of the first data; and   executing training of the machine learning model based on the first data and the label.   
     
     
         2 . The non-transitory recording medium according to  claim 1 , wherein the independence is a mutual information amount between the prediction result and the value of the first attribute. 
     
     
         3 . The non-transitory recording medium to  claim 1 , wherein selecting the first data includes selecting based on a degree of improvement in prediction accuracy of the machine learning model due to each of the plurality of items of data based on an uncertainty of a prediction result in a case in which each of the plurality of items of data is input to the machine learning model, and based on the independence. 
     
     
         4 . The non-transitory recording medium according to  claim 3 , wherein the degree of improvement increases the closer the first data is to a decision boundary of the machine learning model. 
     
     
         5 . The non-transitory recording medium according to  claim 3 , wherein selecting the first data includes selecting a predetermined number of items of the data in which an index represented by the degree of improvement, the independence, and a coefficient representing a trade-off between the degree of improvement and the independence, is equal to or more than a predetermined value, or selecting a predetermined number of items of the data having the highest indices. 
     
     
         6 . The non-transitory recording medium to  claim 1 , wherein:
 calculating the independence includes configuring a part of the plurality of items of data as verification data and configuring data other than the verification data among the plurality of items of data as candidate data, and calculating an independence between a prediction result in a case in which each item of the verification data is input to the machine learning model, and the value of the first attribute of each item of the verification data, the independence being predicated on an independence between a prediction result in a case in which each item of the candidate data is input to the machine learning model, and the value of the first attribute of each item of the candidate data, and   selecting the first data includes selecting the first data from the candidate data.   
     
     
         7 . A machine learning method executable by a computer to perform a process, the process comprising:
 calculating an independence between a prediction result in a case in which each of a plurality of items of unlabeled data is input to a machine learning model, and a value of a first attribute of each of the plurality of items of data;   selecting first data from the plurality of items of data based on the independence;   acquiring a label of the first data; and   executing training of the machine learning model based on the first data and the label.   
     
     
         8 . The machine learning method according to  claim 7 , wherein the independence is a mutual information amount between the prediction result and the value of the first attribute. 
     
     
         9 . The machine learning method according to  claim 7 , wherein selecting the first data includes selecting based on a degree of improvement in prediction accuracy of the machine learning model due to each of the plurality of items of data based on an uncertainty of a prediction result in a case in which each of the plurality of items of data is input to the machine learning model, and based on the independence. 
     
     
         10 . The machine learning method according to  claim 9 , wherein the degree of improvement increases the closer the first data is to a decision boundary of the machine learning model. 
     
     
         11 . The machine learning method according to  claim 9 , wherein selecting the first data includes selecting a predetermined number of items of the data in which an index represented by the degree of improvement, the independence, and a coefficient representing a trade-off between the degree of improvement and the independence, is equal to or more than a predetermined value, or selecting a predetermined number of items of the data having the highest indices. 
     
     
         12 . The machine learning method according to  claim 7 , wherein:
 calculating the independence includes configuring a part of the plurality of items of data as verification data and configuring data other than the verification data among the plurality of items of data as candidate data, and calculating an independence between a prediction result in a case in which each item of the verification data is input to the machine learning model, and the value of the first attribute of each item of the verification data, the independence being predicated on an independence between a prediction result in a case in which each item of the candidate data is input to the machine learning model, and the value of the first attribute of each item of the candidate data, and   selecting the first data includes selecting the first data from the candidate data.   
     
     
         13 . A machine learning apparatus, comprising:
 a memory; and   a processor coupled to the memory, the processor being configured to execute processing, the processing including:   calculating an independence between a prediction result in a case in which each of a plurality of items of unlabeled data is input to a machine learning model, and a value of a first attribute of each of the plurality of items of data;   selecting first data from the plurality of items of data based on the independence;   acquiring a label of the first data; and   executing training of the machine learning model based on the first data and the label.   
     
     
         14 . The machine learning apparatus according to  claim 13 , wherein the independence is a mutual information amount between the prediction result and the value of the first attribute. 
     
     
         15 . The machine learning apparatus according to  claim 13 , wherein selecting the first data includes selecting based on a degree of improvement in prediction accuracy of the machine learning model due to each of the plurality of items of data based on an uncertainty of a prediction result in a case in which each of the plurality of items of data is input to the machine learning model, and based on the independence. 
     
     
         16 . The machine learning apparatus according to  claim 15 , wherein the degree of improvement increases the closer the first data is to a decision boundary of the machine learning model. 
     
     
         17 . The machine learning apparatus according to  claim 15 , wherein selecting the first data includes selecting a predetermined number of items of the data in which an index represented by the degree of improvement, the independence, and a coefficient representing a trade-off between the degree of improvement and the independence, is equal to or more than a predetermined value, or selecting a predetermined number of items of the data having the highest indices. 
     
     
         18 . The machine learning apparatus according to  claim 13 , wherein:
 calculating the independence includes configuring a part of the plurality of items of data as verification data and configuring data other than the verification data among the plurality of items of data as candidate data, and calculating an independence between a prediction result in a case in which each item of the verification data is input to the machine learning model, and the value of the first attribute of each item of the verification data, the independence being conditioned on an independence between a prediction result in a case in which each item of the candidate data is input to the machine learning model, and the value of the first attribute of each item of the candidate data, and   selecting the first data includes selecting the first data from the candidate data.

Join the waitlist — get patent alerts

Track US2025307714A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.