Systems and methods for transfer learning of neural networks
Abstract
Methods and systems may be used for transfer learning of neural networks. According to one example, a method includes: grouping data objects of a first training set into a plurality of clusters; training a base model using a first cluster of the plurality of clusters, the base model being a neural network having a plurality of nodes; generalizing the base model to obtain a generalized base model, the generalizing the base model including setting a portion of the plurality of nodes to have random or predetermined weights; determining that the first cluster is, out of the plurality clusters, most similar to a second training set; and training the generalized base model using the second training set to obtain a trained model.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A computer-implemented method for transfer learning of neural networks, the method comprising:
grouping data objects of a first training set into a plurality of clusters; training a base model using a first cluster of the plurality of clusters, the base model being a neural network having a plurality of nodes; determining a performance of the base model without a random subset of existing nodes of the plurality of nodes of the base model; generalizing the base model to obtain a generalized base model, wherein generalizing the base model includes:
setting a portion of each of the plurality of nodes to have a predetermined weight based on the determined performance of the base model; and
generalizing the base model based on the set portion of each of the plurality of nodes to have the predetermined weight;
determining that the first cluster is, out of the plurality clusters, most similar to a second training set; and training the generalized base model using the second training set to obtain a trained model.
22 . The computer-implemented method of claim 21 , wherein the predetermined weight is set based on determining that removal of the random subset of existing nodes does not substantially affect accuracy of the base model.
23 . The computer-implemented method of claim 21 , wherein the predetermined weight is set for each of the plurality of nodes.
24 . The computer-implemented method of claim 21 , wherein the predetermined weight is set for at least one of the plurality of nodes.
25 . The computer-implemented method of claim 21 , wherein the predetermined weight is zero.
26 . The computer-implemented method of claim 21 , wherein the base model is further trained using the predetermined weight for the plurality of nodes.
27 . The computer-implemented method of claim 21 , wherein the predetermined weight is set by:
pruning the base model to remove one or more nodes of the plurality of nodes from the base model; restoring to the base model the one or more nodes of the plurality of nodes from the base model; and setting the one or more nodes restored in the restoring to have a predetermined weight.
28 . The computer-implemented method of claim 27 , wherein each of the one or more restored nodes has an output.
29 . The computer-implemented method of claim 28 , wherein the output of each of the one or more restored nodes has a predetermined weight.
30 . The computer-implemented method of claim 21 , further comprising:
setting a portion of each of the plurality of nodes to have a random weight based on the determined performance of the base model.
31 . A computer system for transfer learning of neural networks, the computer system comprising:
a memory storing instructions; and one or more processors configured to execute the instructions to perform operations including:
grouping data objects of a first training set into a plurality of clusters;
training a base model using a first cluster of the plurality of clusters, the base model being a neural network having a plurality of nodes;
determining a performance of the base model without a random subset of existing nodes of the plurality of nodes of the base model;
generalizing the base model to obtain a generalized base model, wherein generalizing the base model includes:
setting a portion of each of the plurality of nodes to have a predetermined weight based on the determined performance of the base model; and
generalizing the base model based on the set portion of each of the plurality of nodes to have the predetermined weight;
determining that the first cluster is, out of the plurality clusters, most similar to a second training set; and
training the generalized base model using the second training set to obtain a trained model.
32 . The computer system of claim 31 , wherein the predetermined weight is set based on determining that removal of the random subset of existing nodes does not substantially affect accuracy of the base model.
33 . The computer system of claim 31 , wherein the predetermined weight is set for at least one of the plurality of nodes.
34 . The computer system of claim 31 , wherein the predetermined weight is zero.
35 . The computer system of claim 31 , wherein the base model is further trained using the predetermined weight for the plurality of nodes.
36 . The computer system of claim 31 , wherein the predetermined weight is set by:
pruning the base model to remove one or more nodes of the plurality of nodes from the base model; restoring to the base model the one or more nodes of the plurality of nodes from the base model; and setting the one or more nodes restored in the restoring to have a predetermined weight.
37 . The computer system of claim 36 , wherein each of the one or more restored nodes has an output.
38 . The computer system of claim 37 , wherein the output of each of the one or more restored nodes has a predetermined weight.
39 . The computer system of claim 31 , further comprising:
setting a portion of each of the plurality of nodes to have a random weight based on the determined performance of the base model.
40 . A computer-implemented method for transfer learning of neural networks, the method comprising:
grouping data objects of a first training set into a plurality of clusters; training a base model, the base model being a neural network having a plurality of nodes, using a first cluster of the plurality of clusters and a predetermined weight for the plurality of nodes; determining a performance of the base model without a random subset of existing nodes of the plurality of nodes of the base model; generalizing the base model to obtain a generalized base model, wherein generalizing the base model includes:
setting a portion of each of the plurality of nodes to have the predetermined weight based on the determined performance of the base model, wherein the predetermined weight is set based on determining that removal of the nodes does not substantially affect accuracy of the base model; and
generalizing the base model based on the set portion of each of the plurality of nodes to have the predetermined weight;
determining that the first cluster is, out of the plurality clusters, most similar to a second training set; and training the generalized base model using the second training set to obtain a trained model.Join the waitlist — get patent alerts
Track US2023394372A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.