Neural Architecture Search with Factorized Hierarchical Search Space
Abstract
The present disclosure is directed to an automated neural architecture search approach for designing new neural network architectures such as, for example, resource-constrained mobile CNN models. In particular, the present disclosure provides systems and methods to perform neural architecture search using a novel factorized hierarchical search space that permits layer diversity throughout the network, thereby striking the right balance between flexibility and search space size. The resulting neural architectures are able to be run relatively faster and using relatively fewer computing resources (e.g., less processing power, less memory usage, less power consumption, etc.), all while remaining competitive with or even exceeding the performance (e.g., accuracy) of current state-of-the-art mobile-optimized models.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A computing system comprising:
one or more processors; and one or more non-transitory computer-readable media that store:
a convolutional neural network configured to process an input image to generate a prediction, the convolutional neural network comprising:
an initial convolutional layer configured to receive and process the input image to generate a first intermediate representation;
a plurality of inverted residual bottleneck blocks arranged in a sequence one after another, the plurality of inverted residual bottleneck blocks configured to receive and process the first intermediate representation to generate a second intermediate representation, each of the plurality of inverted residual bottleneck blocks comprising one or more layer reps, each layer rep comprising:
a convolutional layer configured to apply a depthwise convolution;
a convolutional layer configured to apply a pointwise convolution; and
a linear bottleneck layer; and
one or more subsequent layers configured to receive and process the second intermediate representation to generate the prediction; and
instructions that, when executed by the one or more processors, cause the computing system to process the input image with the convolutional neural network to generate the prediction.
22 . The computing system of claim 21 , wherein the initial convolutional layer applies a three-by-three filter.
23 . The computing system of claim 21 , wherein the initial convolutional layer comprises thirty-two channels.
24 . The computing system of claim 21 , wherein the plurality of inverted residual bottleneck blocks comprises seven inverted residual bottleneck blocks.
25 . The computing system of claim 24 , wherein each of at least a first sequential inverted residual bottleneck block, a second sequential inverted residual bottleneck block, and a fourth sequential inverted residual bottleneck block of the seven inverted residual bottleneck blocks apply three-by-three filters.
26 . The computing system of claim 24 , wherein a first sequential inverted residual bottleneck block of the seven inverted residual bottleneck blocks has an expansion factor of one and each of a second, third, fourth, fifth, sixth, and seventh sequential inverted residual bottleneck block of the seven inverted residual bottleneck blocks has an expansion factor of six.
27 . The computing system of claim 21 , wherein a first sequential inverted residual bottleneck block of the plurality of inverted residual bottleneck blocks comprises sixteen channels.
28 . The computing system of claim 21 , wherein a second sequential inverted residual bottleneck block of the plurality of inverted residual bottleneck blocks comprises twenty-four channels.
29 . The computing system of claim 21 , wherein a final sequential inverted residual bottleneck block of the plurality of inverted residual bottleneck blocks comprises three-hundred and twenty channels.
30 . The computing system of claim 21 , wherein a first sequential inverted residual bottleneck block of the plurality of inverted residual bottleneck blocks comprises a single layer rep.
31 . The computing system of claim 21 , wherein a second sequential inverted residual bottleneck block of the plurality of inverted residual bottleneck blocks comprises two layer reps.
32 . The computing system of claim 21 , wherein a final sequential inverted residual bottleneck block of the plurality of inverted residual bottleneck blocks comprises a single layer rep.
33 . The computing system of claim 21 , wherein a first sequential inverted residual bottleneck block of the plurality of inverted residual bottleneck blocks has two-thirds has many channels as a second sequential inverted residual bottleneck block of the plurality of inverted residual bottleneck blocks.
34 . The computing system of claim 21 , wherein a first sequential inverted residual bottleneck block of the plurality of inverted residual bottleneck blocks has one-half has many layer reps as a second sequential inverted residual bottleneck block of the plurality of inverted residual bottleneck blocks.
35 . The computing system of claim 21 , wherein the one or more subsequent layers perform at least one subsequent convolution and a pooling operation.
36 . One or more non-transitory computer-readable media that store:
a convolutional neural network configured to process an input image to generate a prediction, the convolutional neural network comprising:
an initial convolutional layer configured to receive and process the input image to generate a first intermediate representation;
a plurality of inverted residual bottleneck blocks arranged in a sequence one after another, the plurality of inverted residual bottleneck blocks configured to receive and process the first intermediate representation to generate a second intermediate representation, each of the plurality of inverted residual bottleneck blocks comprising a linear layer and depthwise separable convolutional layers; and
one or more subsequent layers configured to receive and process the second intermediate representation to generate the prediction; and
instructions that, when executed by the one or more processors, cause the computing system to process the input image with the convolutional neural network to generate the prediction.
37 . The one or more non-transitory computer-readable media of claim 36 , wherein the initial convolutional layer comprises thirty-two channels.
38 . The one or more non-transitory computer-readable media of claim 36 , wherein the plurality of inverted residual bottleneck blocks comprises seven inverted residual bottleneck blocks.
39 . The one or more non-transitory computer-readable media of claim 38 , wherein each of at least a first sequential inverted residual bottleneck block, a second sequential inverted residual bottleneck block, and a fourth sequential inverted residual bottleneck block of the seven inverted residual bottleneck blocks apply three-by-three filters.
40 . The one or more non-transitory computer-readable media of claim 38 , wherein a first sequential inverted residual bottleneck block of the seven inverted residual bottleneck blocks has an expansion factor of one and each of a second, third, fourth, fifth, sixth, and seventh sequential inverted residual bottleneck block of the seven inverted residual bottleneck blocks has an expansion factor of six.Join the waitlist — get patent alerts
Track US2022101090A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.