US2022414425A1PendingUtilityA1

Resource constrained neural network architecture search

Assignee: GOOGLE LLCPriority: Aug 23, 2019Filed: Aug 19, 2022Published: Dec 29, 2022
Est. expiryAug 23, 2039(~13.1 yrs left)· nominal 20-yr term from priority
G06N 3/063G06F 18/2415G06F 16/9024G06N 3/082G06N 3/04G06F 18/241G06N 20/00G06N 3/044G06N 3/045G06N 3/047G06N 3/09G06N 3/0985G06N 3/0464
67
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, and systems, including computer programs encoded on computer storage media for neural network architecture search. A method includes defining a neural network computational cell, the computational cell including a directed graph of nodes representing respective neural network latent representations and edges representing respective operations that transform a respective neural network latent representation; replacing each operation that transforms a respective neural network latent representation with a respective linear combination of candidate operations, where each candidate operation in a respective linear combination has a respective mixing weight that is parameterized by one or more computational cell hyper parameters; iteratively adjusting values of the computational cell hyper parameters and weights to optimize a validation loss function subject to computational resource constraints; and generating a neural network for performing a machine learning task using the defined computational cell and the adjusted values of the computational cell hyper parameters and weights.

Claims

exact text as granted — not AI-modified
1 . A computer-implemented method when executed by data processing hardware of a user device causes the data processing hardware to perform operations comprising:
 defining a plurality of computational cells of a neural network, each computational cell of the plurality of computational cells comprising a different directed graph of a predetermined number of nodes and edges and one or more respective computational cell hyper parameters, each node representing a respective neural network latent representation and each edge representing a respective operation that transforms a respective neural network latent representation;   for each computational cell of the plurality of computational cells, optimizing a validation loss function subject to one or more computational resource constraints; and   based on each optimized validation loss function, generating the neural network for performing a machine learning task using the respective one or more computational cell hyper parameters of each of the computational cells in the plurality of computational cells.   
     
     
         2 . The method of  claim 1 , wherein the operations further comprise, for each computational cell of the plurality of computational cells, replacing each respective operation that transforms a respective neural network latent representation with a respective linear combination of candidate operations from a predefined set of candidate operations, each candidate operation in a respective linear combination having a respective mixing weight that is parameterized by the respective one or more computational cell hyper parameters before optimizing the validation loss function subject to the one or more computational resource constraints. 
     
     
         3 . The method of  claim 1 , wherein optimizing the validation loss function subject to the one or more computational resource constraints further comprises iteratively adjusting values of the respective one or more computational cell hyper parameters and computational cell weights. 
     
     
         4 . The method of  claim 3 , wherein iteratively adjusting the values of the respective one or more computational cell hyper parameters and the computational cell weights comprises performing a bi-level optimization of the validation loss function and a training loss function that represents a measure of error obtained on training data, wherein the respective one or more computational cell hyper parameters comprise upper level parameters and the computational cell weights comprise lower level parameters. 
     
     
         5 . The method of  claim 3 , wherein iteratively adjusting the values of the computational cell hyper parameters and the computational cell weights comprises defining a respective cost function for each computational resource constraint, each defined cost function mapping the computational cell hyper parameters to a respective resource cost. 
     
     
         6 . The method of  claim 5 , wherein a respective resource cost of an edge in each computational cell is calculated as a softmax over costs of operations in a candidate set of operations. 
     
     
         7 . The method of  claim 5 , wherein the operations further comprise setting lower and higher bound constraints for each defined cost function. 
     
     
         8 . The method of  claim 1 , wherein the validation loss function represents a measure of error obtained after running a validation dataset through each defined computational cell of the plurality of computational cells. 
     
     
         9 . The method of  claim 1 , wherein the one or more computational resource constraints comprise user defined constraints on one or more of memory, number of float point operations, or inference speed. 
     
     
         10 . The method of  claim 1 , wherein the operations further comprise:
 training the generated neural network on training data to obtain a trained neural network; and   performing the machine learning task using the trained neural network.   
     
     
         11 . A system comprising:
 data processing hardware; and   memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
 defining a plurality of computational cells of a neural network, each computational cell of the plurality of computational cells comprising a different directed graph of a predetermined number of nodes and edges and respective one or more computational cell hyper parameters, each node representing a respective neural network latent representation and each edge representing a respective operation that transforms a respective neural network latent representation; 
 for each computational cell of the plurality of computational cells, optimizing a validation loss function subject to one or more computational resource constraints; and 
 based on each optimized validation loss function, generating the neural network for performing a machine learning task using the respective one or more computational cell hyper parameters of each of the computational cells in the plurality of computational cells. 
   
     
     
         12 . The system of  claim 11 , wherein the operations further comprise, for each computational cell of the plurality of computational cells, replacing each respective operation that transforms a respective neural network latent representation with a respective linear combination of candidate operations from a predefined set of candidate operations, each candidate operation in a respective linear combination having a respective mixing weight that is parameterized by the respective one or more computational cell hyper parameters before optimizing the validation loss function subject to the one or more computational resource constraints. 
     
     
         13 . The system of  claim 11 , wherein optimizing the validation loss function subject to the one or more computational resource constraints further comprises iteratively adjusting values of the respective one or more computational cell hyper parameters and computational cell weights. 
     
     
         14 . The system of  claim 13 , wherein iteratively adjusting the values of the respective one or more computational cell hyper parameters and the computational cell weights comprises performing a bi-level optimization of the validation loss function and a training loss function that represents a measure of error obtained on training data, wherein the respective one or more computational cell hyper parameters comprise upper level parameters and the computational cell weights comprise lower level parameters. 
     
     
         15 . The system of  claim 13 , wherein iteratively adjusting the values of the computational cell hyper parameters and the computational cell weights comprises defining a respective cost function for each computational resource constraint, each defined cost function mapping the computational cell hyper parameters to a respective resource cost. 
     
     
         16 . The system of  claim 15 , wherein a respective resource cost of an edge in each computational cell is calculated as a softmax over costs of operations in a candidate set of operations. 
     
     
         17 . The system of  claim 15 , wherein the operations further comprise setting lower and higher bound constraints for each defined cost function. 
     
     
         18 . The system of  claim 11 , wherein the validation loss function represents a measure of error obtained after running a validation dataset through each defined computational cell of the plurality of computational cells. 
     
     
         19 . The system of  claim 11 , wherein the one or more computational resource constraints comprise user defined constraints on one or more of memory, number of float point operations, or inference speed. 
     
     
         20 . The system of  claim 11 , wherein the operations further comprise:
 training the generated neural network on training data to obtain a trained neural network; and   performing the machine learning task using the trained neural network.

Join the waitlist — get patent alerts

Track US2022414425A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.