Systems and methods for simultaneous network pruning and parameter optimization
Abstract
Systems and methods for simultaneous network pruning and parameter optimization are disclosed. A method may include: (1) receiving a network to optimize, the network comprising a plurality of layers; (2) selecting layer pruning and/or channel pruning for the layers within the network; (3) providing a gating module at layers within the network, wherein each gating module opens or closes a gate in the gating module based on an output of a binary head; (4) training parameters for the network; (5) extracting gate open/close features from the gating modules; (6) optimizing a loss function for the network using a polarization regularizer to reach a consensus static sub-network; and (7) updating parameters for each gating module in the network consistent with the consensus static sub-network.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for network pruning, comprising:
receiving, by a network optimization computer program, a network to optimize, the network comprising a plurality of layers; selecting, by the network optimization computer program, layer pruning and/or channel pruning for the layers within the network; providing, by the network optimization computer program, a gating module at layers within the network, wherein each gating module opens or closes a gate in the gating module based on an output of a binary head; training, by the network optimization computer program, parameters for the network; extracting, by the network optimization computer program, gate open/close features from the gating modules; optimizing, by the network optimization computer program, a loss function for the network, wherein the network optimization computer program uses a polarization regularizer to reach a consensus static sub-network; and updating, by the network optimization computer program, parameters for each gating module in the network consistent with the consensus static sub-network.
2 . The method of claim 1 , wherein the gating modules comprise a channel pruning gating module and/or a layer pruning gating module.
3 . The method of claim 1 , wherein the network comprises a residual network.
4 . The method of claim 1 , wherein the network comprises a sequential network.
5 . The method of claim 1 , further comprising:
receiving, by the network optimization computer program, a sparsity hyperparameter.
6 . The method of claim 1 , wherein the gating modules are added at a beginning of each layer and each channel within the network.
7 . The method of claim 1 , wherein each gating module comprises a fully connected layer and a binary head, wherein the fully connected layer that receives a one-dimensional vector, multiplies the one-dimensional vector by a weight matrices, and outputs an output vector, and the binary head receives the output vector and returns a binary value indicating whether the layer will be computed.
8 . The method of claim 7 , wherein the binary head comprises a straight-through estimator.
9 . The method of claim 8 , wherein a gradient of the straight-through estimator updates the parameters of the gating modules through back propagation.
10 . The method of claim 9 , wherein the parameters comprise weight matrices.
11 . A non-transitory computer readable storage medium, including instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to perform steps comprising:
receiving a network to optimize, the network comprising a plurality of layers; selecting layer pruning and/or channel pruning for the layers within the network; providing a gating module at layers within the network, wherein each gating module opens or closes a gate in the gating module based on an output of a binary head; training parameters for the network; extracting gate open/close features from the gating modules; optimizing a loss function for the network using a polarization regularizer to reach a consensus static sub-network; and updating parameters for each gating module in the network consistent with the consensus static sub-network.
12 . The non-transitory computer readable storage medium of claim 11 , wherein the gating modules comprise a channel pruning gating module and/or a layer pruning gating module.
13 . The non-transitory computer readable storage medium of claim 11 , wherein the network comprises a residual network.
14 . The non-transitory computer readable storage medium of claim 11 , wherein the network comprises a sequential network.
15 . The non-transitory computer readable storage medium of claim 11 , further comprising instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to receive a sparsity hyperparameter.
16 . The non-transitory computer readable storage medium of claim 11 , wherein the gating modules are added a beginning of each layer and each channel within the network.
17 . The non-transitory computer readable storage medium of claim 11 , wherein each gating module comprises a fully connected layer and a binary head, wherein the fully connected layer that receives a one-dimensional vector, multiplies the one-dimensional vector by a weight matrices, and outputs an output vector, and the binary head receives the output vector and returns a binary value indicating whether the layer will be computed.
18 . The non-transitory computer readable storage medium of claim 17 , wherein the binary head comprises a straight-through estimator.
19 . The non-transitory computer readable storage medium of claim 18 , wherein a gradient of the straight-through estimator updates the parameters of the gating modules through back propagation.
20 . The non-transitory computer readable storage medium of claim 19 , wherein the parameters comprise weight matrices.Join the waitlist — get patent alerts
Track US2024054346A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.