Apparatuses, methods, and systems for neural networks
Abstract
Methods and apparatuses relating to processing neural networks are described. In one embodiment, an apparatus to process a neural network includes a plurality of fully connected layer chips coupled by an interconnect; a plurality of convolutional layer chips each coupled by an interconnect to a respective fully connected layer chip of the plurality of fully connected layer chips and each of the plurality of fully connected layer chips and the plurality of convolutional layer chips including an interconnect to couple each of a forward propagation compute intensive tile, a back propagation compute intensive tile, and a weight gradient compute intensive tile of a column of compute intensive tiles between a first memory intensive tile and a second memory intensive tile.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An apparatus comprising:
a first parallel compute device comprising a plurality of chips, the plurality of chips including a first chip coupled to a first memory device and a second chip coupled to a second memory device; a first interconnect to couple the first chip and the second chip; and an interface to couple the first parallel compute device to one or more additional parallel compute devices; the first chip comprising:
a first scalar register file to store data related to control flow operations;
first scalar execution circuitry to execute one or more scalar instructions to perform control flow operations in accordance with the data to control execution of an instruction sequence, the control flow operations including program loops, branches, and address calculations, and the instruction sequence including a first matrix multiply instruction;
a first decoder to decode the instructions of the instruction sequence including the first matrix multiply instruction, the first matrix multiply instruction to specify multiplication of a first matrix by a second matrix; and
a first matrix processing unit comprising a first array of processing elements to perform a plurality of parallel fused multiply-accumulate operations in accordance with the first matrix multiply instruction to multiply first data elements of the first matrix by corresponding second data elements of the second matrix to generate a corresponding plurality of products, and to add the plurality of products to corresponding accumulated values to generate a result matrix.
2 . The apparatus of claim 1 , wherein the first matrix multiply instruction is to further specify a first size associated with the first matrix and a second size associated with the second matrix.
3 . The apparatus of claim 1 , wherein the second chip comprises:
a second scalar register file to store data related to control flow of a first program or a second program; second scalar execution circuitry to execute one or more scalar instructions to perform control flow operations related to execution of the second program in accordance with the data, the control flow operations including program loops, branches, and address calculations; a second decoder to decode a second matrix multiply instruction, the second matrix multiply instruction to specify multiplication of a third matrix by a fourth matrix, the second matrix multiply instruction to indicate a third matrix size of the third matrix and a fourth matrix size of the fourth matrix; and a second matrix processing unit comprising a second array of processing elements to perform a second plurality of parallel fused multiply-accumulate operations in accordance with the second matrix multiply instruction to multiply third data elements of the third matrix by corresponding fourth data elements of the fourth matrix to generate a second corresponding plurality of products, and to add the second corresponding plurality of products to second corresponding accumulated values to generate a second result matrix.
4 . The apparatus of claim 1 , wherein the first data elements comprise convolution input elements.
5 . The apparatus of claim 4 , wherein the second data elements comprise neural network weights.
6 . The apparatus of claim 1 , wherein the first matrix multiply instruction comprises a first operand to identify the first data elements and a second operand to identify the second data elements.
7 . The apparatus of claim 6 , wherein the first operand identifies the first data elements in a first one or more registers and the second operand identifies the second data elements in a second one or more registers.
8 . The apparatus claim 1 , further comprising:
an instruction fetch unit to fetch the first matrix multiply instruction; a decoder to decode the first matrix multiply instruction to generate parallel multiply-add operations; and a scheduler to schedule the parallel multiply-add operations for execution by at least a portion of the array of processing elements.
9 . The apparatus of claim 1 , wherein the first memory device and the second memory device comprise high bandwidth memory (HBM) devices.
10 . The apparatus of claim 3 , wherein each of the first parallel compute device and the one or more additional parallel compute devices are to access a system memory using a shared address range.
11 . The apparatus of claim 10 , further comprising coherency logic to ensure coherency of data in the system memory which is shared between the first parallel compute device and the one or more additional parallel compute devices.
12 . The apparatus of claim 11 , wherein the first program comprises one or more machine learning tasks, wherein separate portions of the one or more machine learning tasks are to be executed by the first parallel compute device and the one or more additional parallel compute devices.
13 . An apparatus comprising:
a compute processor package including a first die and a second die, the first die coupled to a first memory device and the second die coupled to a second memory device; a first interconnect to couple the first die and the second die; and an interface to couple the compute processor package to one or more additional compute processor packages; at least the first die comprising: a first scalar register file to store data related to control flow operations; first scalar execution circuitry to execute one or more scalar instructions to perform control flow operations in accordance with the data to control execution of an instruction sequence, the control flow operations including program loops, branches, and address calculations, and the instruction sequence including a first matrix multiply instruction; a first decoder to decode the instructions of the instruction sequence including the first matrix multiply instruction, the first matrix multiply instruction to specify multiplication of a first matrix by a second matrix, and to specify a first size associated with the first matrix and a second size associated with the second matrix; and a first matrix processing unit comprising a first array of processing elements to perform a plurality of parallel fused multiply-accumulate operations in accordance with the first matrix multiply instruction to multiply first data elements of the first matrix by corresponding second data elements of the second matrix to generate a corresponding plurality of products, and to add the plurality of products to corresponding accumulated values to generate a result matrix.
14 . The apparatus of claim 13 , wherein the second die comprises:
a second scalar register file to store data related to control flow of a first program or a second program; second scalar execution circuitry to execute one or more scalar instructions to perform control flow operations related to execution of the second program in accordance with the data, the control flow operations including program loops, branches, and address calculations; a second decoder to decode a second matrix multiply instruction, the second matrix multiply instruction to specify multiplication of a third matrix by a fourth matrix, the second matrix multiply instruction to indicate a third matrix size of the third matrix and a fourth matrix size of the fourth matrix; and a second matrix processing unit comprising a second array of processing elements to perform a second plurality of parallel fused multiply-accumulate operations in accordance with the second matrix multiply instruction to multiply third data elements of the third matrix by corresponding fourth data elements of the fourth matrix to generate a second corresponding plurality of products, and to add the second corresponding plurality of products to second corresponding accumulated values to generate a second result matrix.
15 . The apparatus of claim 13 , wherein the first data elements comprise convolution input elements.
16 . The apparatus of claim 15 , wherein the second data elements comprise neural network weights.
17 . The apparatus of claim 13 , wherein the first matrix multiply instruction comprises a first operand to identify the first data elements and a second operand to identify the second data elements.
18 . The apparatus of claim 17 , wherein the first operand identifies the first data elements in a first one or more registers and the second operand identifies the second data elements in a second one or more registers.
19 . The apparatus claim 13 , further comprising:
an instruction fetch unit to fetch the first matrix multiply instruction; a decoder to decode the first matrix multiply instruction to generate parallel multiply-add operations; and a scheduler to schedule the parallel multiply-add operations for execution by at least a portion of the array of processing elements.
20 . The apparatus of claim 13 , wherein the first memory device and the second memory device comprise high bandwidth memory (HBM) devices.
21 . The apparatus of claim 14 , wherein each of the first matrix processing unit and the second matrix processing unit are to access a system memory using a shared address range.
22 . The apparatus of claim 21 , further comprising coherency logic to ensure coherency of data in the system memory which is shared between the first matrix processing unit and the second matrix processing unit.
23 . The apparatus of claim 22 , wherein the first program comprises one or more machine learning tasks, wherein separate portions of the one or more machine learning tasks are to be executed by the first matrix processing unit and the second matrix processing unit.Join the waitlist — get patent alerts
Track US2022050683A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.