US2023359862A1PendingUtilityA1

Systems and Methods for Machine-Learned Models Having Convolution and Attention

Assignee: GOOGLE LLCPriority: May 27, 2021Filed: Jul 19, 2023Published: Nov 9, 2023
Est. expiryMay 27, 2041(~14.8 yrs left)· nominal 20-yr term from priority
G06N 3/0464G06N 3/09G06N 3/044G06N 20/00G06N 3/063G06N 3/08G06V 10/454G06V 10/82G06N 3/084G06N 3/045
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A computer-implemented method for performing computer vision with reduced computational cost and improved accuracy can include obtaining, by a computing system including one or more computing devices, input data comprising an input tensor having one or more dimensions, providing, by the computing system, the input data to a machine-learned convolutional attention network, the machine-learned convolutional attention network including two or more network stages, and, in response to providing the input data to the machine-learned convolutional attention network, receiving, by the computing system, a machine-learning prediction from the machine-learned convolutional attention network. The convolutional attention network can include at least one attention block, wherein the attention block includes a relative attention mechanism, the relative attention mechanism including the sum of a static convolution kernel with an adaptive attention matrix. This provides for improved generalization, capacity, and efficiency of the convolutional attention network relative to some existing models.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 .- 22 . (canceled) 
     
     
         23 . A computer-implemented method comprising:
 obtaining, by a computing system, input data;   providing the input data to a neural network, implemented by the computing system, that includes a statice convolution kernel and an adaptive attention matrix, wherein the neural network processes the input data by applying a sum of a static convolution kernel and the adaptive attention matrix, either prior to or subsequent to performing a Softmax normalization, to the input data; and   receiving, by the computing system, a prediction that is based on the neural network processing the input data.   
     
     
         24 . The computer-implemented method of  claim 23 , wherein the input data comprises an input tensor having one or more dimensions. 
     
     
         25 . The computer-implemented method of  claim 23 , wherein the neural network is a machine-learned convolutional network that includes at least a convolutional stage and an attention stage that comprises a relative attention mechanism that is configured to apply the sum of the static convolution kernel and the adaptive attention matrix, either prior to or subsequent to performing the Softmax normalization, to the input data. 
     
     
         26 . The computer-implemented method of  claim 23 , wherein the neural network includes at least an S 0  stage, an S 1  stage, an S 2  stage, an S 3  stage, and an S 4  stage. 
     
     
         27 . The computer-implemented method of  claim 26 , wherein the S 0  stage comprises a two-layer convolutional stem network. 
     
     
         28 . The computer-implemented method of  claim 26 , wherein the S 1  stage comprises one or more convolutional blocks with squeeze excitation. 
     
     
         29 . The computer-implemented method of  claim 28 , wherein the one or more convolutional blocks of the S 1  stage comprise mobile inverted bottleneck convolution (MBConv) blocks, the MBConv blocks configured to expand channel size from an original channel size of an input to the one or more convolutional blocks and subsequently project the expanded channel size back to the original channel size. 
     
     
         30 . The computer-implemented method of  claim 26 , wherein each of the S 2  stage, S 3  stage, or S 4  stage comprising a convolutional stage comprise a mobile inverted bottleneck convolution (MBConv) block. 
     
     
         31 . The computer-implemented method of  claim 26 , wherein a number of channels is doubled for at least one of the S 1  stage, the S 2  stage, the S 3  stage, or the S 4  stage. 
     
     
         32 . The computer-implemented method of  claim 26 , wherein a width of the S 0  stage is less than or equal to a width of the S 1  stage. 
     
     
         33 . The computer-implemented method of  claim 26 , wherein each of the S 0  stage, the S 1  stage, and the S 4  stage comprises two blocks, and wherein the S 2  stage and the S 3  stage each comprise greater than two blocks. 
     
     
         34 . The computer-implemented method of  claim 23 , wherein a spatial resolution gradually decreases over two or more network stages of the neural network. 
     
     
         35 . The computer-implemented method of  claim 23 , wherein the sum of the static convolution kernel and the adaptive attention matrix is applied to the input data prior to performing the SoftMax normalization . 
     
     
         36 . The computer-implemented method of  claim 23 , wherein the sum of the static convolution kernel and the adaptive attention matrix is applied to the input data subsequent to performing the SoftMax normalization by the relative attention mechanism. 
     
     
         37 . The computer-implemented method of  claim 23 , wherein the prediction comprises a computer vision output. 
     
     
         38 . The computer-implemented method of  claim 23 , wherein the prediction comprises a classification output. 
     
     
         39 . The computer-implemented method of  claim 23 , wherein one or more convolutional stages of the neural network are sequentially prior to one or more attention stages of the neural network. 
     
     
         40 . A computing system comprising:
 one or more processors; and   one or more non-transitory computer-readable media that store instructions that when executed by the one or more processors, cause the computer system to perform operations comprising:   obtaining input data;   providing the input data to a neural network that includes a statice convolution kernel and an adaptive attention matrix, wherein the neural network processes the input data by applying a sum of a static convolution kernel and the adaptive attention matrix, either prior to or subsequent to performing a Softmax normalization, to the input data; and   receiving a prediction from the neural network, wherein the prediction is based on the neural network processing the input data.   
     
     
         41 . The computing system of  claim 40 , wherein the neural network is a machine-learned convolutional network that includes at least a convolutional stage and an attention stage that comprises a relative attention mechanism that is configured to apply the sum of the static convolution kernel and the adaptive attention matrix, either prior to or subsequent to performing the Softmax normalization, to the input data 
     
     
         42 . The computing system of  claim 40 , wherein the neural network includes at least an S 0  stage, an S 1  stage, an S 2  stage, an S 3  stage, and an S 4  stage.

Join the waitlist — get patent alerts

Track US2023359862A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.