US2018197080A1PendingUtilityA1

Learning apparatus and method for bidirectional learning of predictive model based on data sequence

Assignee: IBMPriority: Jan 11, 2017Filed: Jan 11, 2017Published: Jul 12, 2018
Est. expiryJan 11, 2037(~10.5 yrs left)· nominal 20-yr term from priority
G06N 7/01G06N 3/047G06N 3/08G06N 3/0442G06N 3/09G06N 3/0475G06N 5/022
50
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A computer-implemented method and an apparatus are provided for learning a first model. The method includes generating a second model based on the first model. The first model is configured to perform a learning process based on sequentially inputting each of a plurality of pieces of input data that include a plurality of input values and that are from a first input data sequence. The second model is configured to learn a first learning target parameter included in the first model based on inputting, in an order differing from an order in the first model, each of a plurality of pieces of input data that include a plurality of input values and are from a second input data sequence. The method further includes performing a learning process using both the first model and the second model. The method also includes storing the first model that has been learned.

Claims

exact text as granted — not AI-modified
1 . A computer-implemented method for learning a first model, comprising:
 generating, by a processor, a second model based on the first model, the first model being configured to perform a learning process based on sequentially inputting each of a plurality of pieces of input data that include a plurality of input values and that are from a first input data sequence, the second model being configured to learn a first learning target parameter included in the first model based on inputting, in an order differing from an order in the first model, each of a plurality of pieces of input data that include a plurality of input values and are from a second input data sequence;   performing, by the processor, a learning process using both the first model and the second model; and   storing, in a memory device, the first model that has been learned.   
     
     
         2 . The computer-implemented method of  claim 1 , wherein the storing the first model that has been learned includes deleting, from the memory device, the second model that has been learned and outputting the first model that has been learned as a predictive model based on an input data sequence. 
     
     
         3 . The computer-implemented method of  claim 2 , wherein the generating the second model includes generating the second model for learning the learning target parameter by inputting, in a backwards order, each of the plurality of pieces of input data from the second input data sequence. 
     
     
         4 . The computer-implemented method of  claim 3 , wherein the first input data sequence and the second input data sequence are time-series input data sequences, wherein the first model inputs the first input data sequence in order from older to newer ones of the plurality of pieces of input data, and wherein the second model inputs the second input data sequence in order from newer to older ones of the plurality of pieces of input data. 
     
     
         5 . The computer-implemented method of  claim 3 , wherein the first model and the second model each include the first learning target parameter and a second learning target parameter, and wherein the performing the learning process includes:
 learning the second learning target parameter by using the first model without changing the first learning target parameter, and   learning the first learning target parameter by using the second model without changing the second learning target parameter.   
     
     
         6 . The computer-implemented method of  claim 5 , wherein the first learning target parameter is operable to be learned with higher accuracy by learning using the second model than by learning using the first model, and wherein the second learning target parameter is operable to be learned with higher accuracy by learning using the first model than by learning using the second model. 
     
     
         7 . The computer-implemented method of  claim 3 , wherein the first input data sequence and the second input data sequence are at least partially identical. 
     
     
         8 . The computer-implemented method of  claim 3 , wherein the first input data sequence and the second input data sequence are input data sequences for learning that are different from each other and included in a plurality of input data sequences for learning. 
     
     
         9 . The computer-implemented method of  claim 3 , wherein the performing the learning process includes performing the learning process with the first model a greater number of times than the learning process with the second model. 
     
     
         10 . The computer-implemented method of  claim 3 , wherein the performing the learning process includes performing the learning process with the first model using a higher learning rate than is used for the learning process with the second model. 
     
     
         11 . The computer-implemented method of  claim 3 , wherein the performing the learning process includes obtaining the first model that has been learned by performing the learning process with the first model last. 
     
     
         12 . The computer-implemented method of  claim 4 , wherein
 the first model includes a plurality of input nodes that sequentially input a plurality of input values at each time point of the first input data sequence, and a weight parameter between each input node and each input value at a time point before a time point corresponding to the plurality of input nodes, and   the second model includes a plurality of input nodes that input, in a backwards order, a plurality of input values at each time point of the second input data sequence, and a weight parameter between each input node and each input value at a time point after the time point corresponding to the plurality of input nodes.   
     
     
         13 . The computer-implemented method of  claim 12 , wherein
 the first model further includes a weight parameter between each input node and each of a plurality of hidden nodes corresponding to the time point before the time point corresponding to the plurality of input nodes, and a weight parameter between each hidden node and each input value corresponding to the time point before the time point corresponding to the plurality of input nodes, and   the second model further includes a weight parameter between each input node and each of a plurality of hidden nodes corresponding to the time point after the time point corresponding to the plurality of input nodes, and a weight parameter between each hidden node and each input value corresponding to the time point after the time point corresponding to the plurality of input nodes.   
     
     
         14 . The computer-implemented method of  claim 13 , wherein the performing the learning process includes:
 learning the weight parameter between each hidden node and each input value corresponding to the time point before the time point corresponding to the plurality of input nodes in the first mode, using the learning process with the first model; and   learning the weight parameter between each input node and each of the plurality of hidden nodes corresponding to the time point after the time point corresponding to the plurality of input nodes in the second model, using the learning process with the second model.   
     
     
         15 - 20 . (canceled)

Join the waitlist — get patent alerts

Track US2018197080A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.