US2022156553A1PendingUtilityA1

Attention neural networks with sparse attention mechanisms

Assignee: GOOGLE LLCPriority: Jun 5, 2020Filed: Jan 31, 2022Published: May 19, 2022
Est. expiryJun 5, 2040(~13.8 yrs left)· nominal 20-yr term from priority
G06N 3/045G06N 3/096G06N 3/098G06N 3/0455G06N 3/0495G06N 3/09G06N 20/00G06N 3/063G06N 3/084G06N 3/04
57
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for processing network inputs using an attention neural network that has one or more sparse attention sub-layers. Each sparse attention sub-layer is configured to apply a sparse attention mechanism that attends differently for input positions that are in a first proper subset of the input positions in the input to the sub-layer than for positions that are not in the first proper subset.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system for performing a machine learning task on a network input to generate a network output, the system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to implement:
 an attention neural network configured to perform the machine learning task, the attention neural network comprising one or more sparse attention layers, each sparse attention layer comprising one or more sparse attention sub-layers, each sparse attention sub-layer configured to:
 receive a sequence of queries derived from an input sequence to the sparse attention layer, the sequence of queries having a respective query at each of a plurality of input positions; 
 receive a sequence of keys derived from the input sequence to the sparse attention layer, the sequence of keys having a respective key at each of the plurality of input positions; 
 receive a sequence of value inputs derived from the input sequence to the sparse attention layer, the sequence of value inputs having a respective value input at each of the plurality of input positions; and 
 generate an attended input sequence comprising a respective attended input at each of the plurality of input positions, comprising:
 for each input position in a first proper subset of the input positions, generating the attended input at the input position by:
 using the query at the input position to attend over all of the keys in the sequence of keys to generate a respective weight for all of the input positions and computing a weighted sum of the value inputs at all of the input positions in accordance with the respective weights; and 
 
 for each input position in a second proper subset of the input positions, generating the attended input at the input position by:
 using the query at the input position to attend over only the keys at a corresponding proper subset of the input positions to generate a respective weight for each of the input positions in the corresponding proper subset and computing a weighted sum of the value inputs at the corresponding proper subset of the input positions in accordance with the respective weights for the corresponding proper subset of the input positions, the corresponding proper subset of input positions for each input position in the second proper subset including: 
 the first proper subset of the input positions; and 
 one or more input positions outside of the first proper subset of the input positions.

Join the waitlist — get patent alerts

Track US2022156553A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.