Learning apparatus, learning method, and program
Abstract
A learning apparatus according to the present disclosure includes: a first training data generating unit configured to generate first training data in which a vector including values of a plurality of elements output by inputting unlabeled training data to a pre-learned machine learning model is a target variable; a second training data generating unit configured to generate second training data in which the values of the elements of the vector that is the target variable of the first training data are set so that a difference in magnitude of value between at least some of the elements becomes larger; and a learning unit configured to generate a machine learning model by machine learning using the first training data and the second training data.
Claims
exact text as granted — not AI-modified1 . A learning apparatus comprising:
at least one memory configured to store processing instructions; and at least one processor configured to execute the processing instructions to: generate first training data with a vector as a target variable, the vector including values of a plurality of elements output by inputting unlabeled training data to a pre-learned machine learning model; generate second training data in which the values of the elements of the vector as the target variable of the first training data are set so that a difference in magnitude of value between at least some of the elements becomes larger; and generate a machine learning model by machine learning using the first training data and the second training data.
2 . The learning apparatus according to claim 1 , wherein the at least one processor is configured to execute the processing instructions to
generate the second training data in which the values of the elements of the vector as the target variable of the first training data are set so that a difference in value between at least one of the elements that has a large value as compared with others of the elements by a preset criterion and another of the elements becomes larger.
3 . The learning apparatus according to claim 1 , wherein the at least one processor is configured to execute the processing instructions to
generate the second training data in which the values of the elements of the vector as the target variable of the first training data are set so that a value of an element that has a largest value becomes largest and a difference in value between the element and others of the elements becomes larger.
4 . The learning apparatus according to claim 1 , wherein the at least one processor is configured to execute the processing instructions to
generate the second training data by setting, among the values of the elements of the vector as the target variable of the first training data, a value of an element that has a largest value to a value greater than 0 and values of elements other than the element to 0.
5 . The learning apparatus according to claim 1 , wherein the at least one processor is configured to execute the processing instructions to
generate the second training data by setting a value of a temperature parameter of softmax function to a value smaller than 1, the softmax function being used for generation of the vector as the target variable of the first training data.
6 . The learning apparatus according to claim 1 , wherein the at least one processor is configured to execute the processing instructions to
generate the machine learning model by machine learning using the second training data in a preset ratio to the first training data.
7 . The learning apparatus according to claim 6 , wherein the at least one processor is configured to execute the processing instructions to:
calculate a loss function L α by L α =(1−α)L 0 +αL 1 , where a parameter indicating the ratio of the second training data to the first training data is α, a loss function in machine learning using the first training data is L 0 , and a loss function in machine learning using the second training data is L 1 ; and generate the machine learning model based on the loss function L α .
8 . A learning method comprising:
generating first training data with a vector as a target variable, the vector including values of a plurality of elements output by inputting unlabeled training data to a pre-learned machine learning model; generating second training data in which the values of the elements of the vector as the target variable of the first training data are set so that a difference in magnitude of value between at least some of the elements becomes larger; and generating a machine learning model by machine learning using the first training data and the second training data.
9 . The learning method according to claim 8 , comprising
generating the second training data in which the values of the elements of the vector as the target variable of the first training data are set so that a difference in value between at least one of the elements that has a large value as compared with others of the elements by a preset criterion and another of the elements becomes larger.
10 . The learning method according to claim 8 , comprising
generating the second training data in which the values of the elements of the vector as the target variable of the first training data are set so that a value of an element that has a largest value becomes largest and a difference in value between the element and others of the elements becomes larger.
11 . The learning method according to claim 8 , comprising
generating the second training data by setting, among the values of the elements of the vector as the target variable of the first training data, a value of an element that has a largest value to a value greater than 0 and values of elements other than the element to 0.
12 . The learning method according to claim 8 , comprising
generating the second training data by setting a value of a temperature parameter of softmax function to a value smaller than 1, the softmax function being used for generation of the vector as the target variable of the first training data.
13 . The learning method according to claim 8 , comprising
generating the machine learning model by machine learning using the second training data in a preset ratio to the first training data.
14 . The learning method according to claim 13 , comprising:
calculating a loss function L α by L α =(1−α)L 0 +αL 1 , where α is a parameter indicating the ratio of the second training data to the first training data, L 0 is a loss function in machine learning using the first training data, and L 1 is a loss function in machine learning using the second training data; and generating the machine learning model based on the loss function L α .
15 . A non-transitory computer-readable storage medium storing a program, the program comprising instructions for causing a computer to execute processes to:
generate first training data with a vector as a target variable, the vector including values of a plurality of elements output by inputting unlabeled training data to a pre-learned machine learning model; generate second training data in which the values of the elements of the vector as the target variable of the first training data are set so that a difference in magnitude of value between at least some of the elements becomes larger; and generate a machine learning model by machine learning using the first training data and the second training data.Join the waitlist — get patent alerts
Track US2024211808A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.