US2022121954A1PendingUtilityA1

Distributed convolution for neural networks

Assignee: INTEL CORPPriority: Dec 30, 2016Filed: Dec 28, 2021Published: Apr 21, 2022
Est. expiryDec 30, 2036(~10.4 yrs left)· nominal 20-yr term from priority
G06N 3/045G06N 3/0464G06F 17/153G06N 3/084G06N 3/063G06F 17/16G06N 3/0454
66
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In one embodiment, a matrix operation may be performed using a plurality of input matrices, wherein the matrix operation is associated with one or more convolution operations. The plurality of input matrices may be partitioned into a plurality of input partitions, wherein the plurality of input matrices is partitioned based on a number of available processing elements. The plurality of input partitions may be distributed among a plurality of processing elements, wherein each input partition is distributed to a particular processing element of the plurality of processing elements. A plurality of partial matrix operations may be performed using the plurality of processing elements, and partial matrix data may be transmitted between the plurality of processing elements while performing the plurality of partial matrix operations. A result of the matrix operation may be determined based on the plurality of partial matrix operations.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method to perform a matrix operation associated with a plurality of input matrices, the method comprising:
 partitioning the plurality of input matrices into a plurality of input partitions, wherein the plurality of input matrices is partitioned based on a number of available processing elements;   distributing the plurality of input partitions among a plurality of processing elements, wherein each input partition is distributed to a particular processing element of the plurality of processing elements;   performing a plurality of partial matrix operations using the plurality of processing elements;   collecting results of the plurality of partial matrix operations; and   determining a result of the matrix operation based on the collected results of the plurality of partial matrix operations.   
     
     
         2 . The method as defined in  claim 1 , wherein the matrix operation is matrix multiplication. 
     
     
         3 . The method as defined in  claim 1 , wherein the performing the plurality of partial matrix operations is performed using one or more field-programmable gate arrays. 
     
     
         4 . The method as defined in  claim 1 , wherein performing a plurality of partial matrix operations includes performing the plurality of partial matrix operations in a plurality of stages. 
     
     
         5 . The method as defined in  claim 1 , wherein the plurality of partial matrix operations are performed via a plurality of matrix processing units (MPUs), wherein each MPU includes processing circuitry to perform matrix arithmetic. 
     
     
         6 . The method as defined in  claim 1 , wherein performing the plurality of partial matrix operations includes performing a first processing stage via a first partition of the input matrices and performing a second processing stage via a second partition of the input matrices. 
     
     
         7 . The method as defined in  claim 1 , wherein the plurality of input matrices includes matrix data associated with one or more images. 
     
     
         8 . At least one non-transitory machine accessible storage medium having instructions stored thereon, the instructions, when executed on a machine, cause the machine to:
 partition a plurality of input matrices into a plurality of input partitions, wherein the plurality of input matrices is partitioned based on a number of available processing elements;   distribute the plurality of input partitions among a plurality of processing elements, wherein each input partition is distributed to a particular processing element of the plurality of processing elements;   perform a plurality of partial matrix operations using the plurality of processing elements;   collect results of the plurality of partial matrix operations; and   determine a result of the matrix operation based on the collected results of the plurality of partial matrix operations.   
     
     
         9 . The at least one non-transitory machine accessible storage medium as defined in  claim 8 , wherein the matrix operation is matrix multiplication. 
     
     
         10 . The at least one non-transitory machine accessible storage medium as defined in  claim 8 , wherein the instructions, when executed, cause the machine to perform the plurality of partial matrix operations using one or more field-programmable gate arrays. 
     
     
         11 . The at least one non-transitory machine accessible storage medium as defined in  claim 8 , wherein the instructions, when executed, cause the machine to perform the plurality of partial matrix operations in a plurality of stages. 
     
     
         12 . The at least one non-transitory machine accessible storage medium as defined in  claim 8 , wherein the instructions, when executed, cause the machine to perform the plurality of partial matrix operations via a plurality of matrix processing units (MPUs), wherein each MPU includes processing circuitry to perform matrix arithmetic. 
     
     
         13 . The at least one non-transitory machine accessible storage medium as defined in  claim 8 , wherein the instructions, when executed, cause the machine to perform the plurality of partial matrix operations via performing a first processing stage via a first partition of the input matrices and performing a second processing stage via a second partition of the input matrices. 
     
     
         14 . The at least one non-transitory machine accessible storage medium as defined in  claim 8 , wherein the plurality of input matrices includes matrix data associated with one or more images. 
     
     
         15 . An apparatus, comprising:
 memory circuitry to store a plurality of input matrices;   processing circuitry to:
 partition the plurality of input matrices into a plurality of input partitions, wherein the plurality of input matrices is partitioned based on a number of available processing elements; 
 distribute the plurality of input partitions among a plurality of processing elements, wherein each input partition is distributed to a particular processing element of the plurality of processing elements; 
 perform a plurality of partial matrix operations using the plurality of processing elements; 
 collect results of the plurality of partial matrix operations; and 
 determine a result of the matrix operation based on the collected results of the plurality of partial matrix operations. 
   
     
     
         16 . The apparatus as defined in  claim 15 , wherein the matrix operation is matrix multiplication. 
     
     
         17 . The apparatus as defined in  claim 15 , wherein the processing circuitry is to perform the plurality of partial matrix operations via one or more field-programmable gate arrays. 
     
     
         18 . The apparatus as defined in  claim 15 , wherein the processing circuitry is to perform the plurality of partial matrix operations in a plurality of stages. 
     
     
         19 . The apparatus as defined in  claim 15 , wherein the processing circuitry is to perform the plurality of partial matrix operations via a plurality of matrix processing units (MPUs), wherein each MPU includes processing circuitry to perform matrix arithmetic. 
     
     
         20 . The apparatus as defined in  claim 15 , wherein the processing circuitry is to perform the plurality of partial matrix operations via performing a first processing stage via a first partition of the input matrices and performing a second processing stage via a second partition of the input matrices. 
     
     
         21 . The apparatus as defined in  claim 15 , wherein the plurality of input matrices includes matrix data associated with one or more images.

Join the waitlist — get patent alerts

Track US2022121954A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.