Distributed convolution for neural networks
Abstract
In one embodiment, a matrix operation may be performed using a plurality of input matrices, wherein the matrix operation is associated with one or more convolution operations. The plurality of input matrices may be partitioned into a plurality of input partitions, wherein the plurality of input matrices is partitioned based on a number of available processing elements. The plurality of input partitions may be distributed among a plurality of processing elements, wherein each input partition is distributed to a particular processing element of the plurality of processing elements. A plurality of partial matrix operations may be performed using the plurality of processing elements, and partial matrix data may be transmitted between the plurality of processing elements while performing the plurality of partial matrix operations. A result of the matrix operation may be determined based on the plurality of partial matrix operations.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method to perform a matrix operation associated with a plurality of input matrices, the method comprising:
partitioning the plurality of input matrices into a plurality of input partitions, wherein the plurality of input matrices is partitioned based on a number of available processing elements; distributing the plurality of input partitions among a plurality of processing elements, wherein each input partition is distributed to a particular processing element of the plurality of processing elements; performing a plurality of partial matrix operations using the plurality of processing elements; collecting results of the plurality of partial matrix operations; and determining a result of the matrix operation based on the collected results of the plurality of partial matrix operations.
2 . The method as defined in claim 1 , wherein the matrix operation is matrix multiplication.
3 . The method as defined in claim 1 , wherein the performing the plurality of partial matrix operations is performed using one or more field-programmable gate arrays.
4 . The method as defined in claim 1 , wherein performing a plurality of partial matrix operations includes performing the plurality of partial matrix operations in a plurality of stages.
5 . The method as defined in claim 1 , wherein the plurality of partial matrix operations are performed via a plurality of matrix processing units (MPUs), wherein each MPU includes processing circuitry to perform matrix arithmetic.
6 . The method as defined in claim 1 , wherein performing the plurality of partial matrix operations includes performing a first processing stage via a first partition of the input matrices and performing a second processing stage via a second partition of the input matrices.
7 . The method as defined in claim 1 , wherein the plurality of input matrices includes matrix data associated with one or more images.
8 . At least one non-transitory machine accessible storage medium having instructions stored thereon, the instructions, when executed on a machine, cause the machine to:
partition a plurality of input matrices into a plurality of input partitions, wherein the plurality of input matrices is partitioned based on a number of available processing elements; distribute the plurality of input partitions among a plurality of processing elements, wherein each input partition is distributed to a particular processing element of the plurality of processing elements; perform a plurality of partial matrix operations using the plurality of processing elements; collect results of the plurality of partial matrix operations; and determine a result of the matrix operation based on the collected results of the plurality of partial matrix operations.
9 . The at least one non-transitory machine accessible storage medium as defined in claim 8 , wherein the matrix operation is matrix multiplication.
10 . The at least one non-transitory machine accessible storage medium as defined in claim 8 , wherein the instructions, when executed, cause the machine to perform the plurality of partial matrix operations using one or more field-programmable gate arrays.
11 . The at least one non-transitory machine accessible storage medium as defined in claim 8 , wherein the instructions, when executed, cause the machine to perform the plurality of partial matrix operations in a plurality of stages.
12 . The at least one non-transitory machine accessible storage medium as defined in claim 8 , wherein the instructions, when executed, cause the machine to perform the plurality of partial matrix operations via a plurality of matrix processing units (MPUs), wherein each MPU includes processing circuitry to perform matrix arithmetic.
13 . The at least one non-transitory machine accessible storage medium as defined in claim 8 , wherein the instructions, when executed, cause the machine to perform the plurality of partial matrix operations via performing a first processing stage via a first partition of the input matrices and performing a second processing stage via a second partition of the input matrices.
14 . The at least one non-transitory machine accessible storage medium as defined in claim 8 , wherein the plurality of input matrices includes matrix data associated with one or more images.
15 . An apparatus, comprising:
memory circuitry to store a plurality of input matrices; processing circuitry to:
partition the plurality of input matrices into a plurality of input partitions, wherein the plurality of input matrices is partitioned based on a number of available processing elements;
distribute the plurality of input partitions among a plurality of processing elements, wherein each input partition is distributed to a particular processing element of the plurality of processing elements;
perform a plurality of partial matrix operations using the plurality of processing elements;
collect results of the plurality of partial matrix operations; and
determine a result of the matrix operation based on the collected results of the plurality of partial matrix operations.
16 . The apparatus as defined in claim 15 , wherein the matrix operation is matrix multiplication.
17 . The apparatus as defined in claim 15 , wherein the processing circuitry is to perform the plurality of partial matrix operations via one or more field-programmable gate arrays.
18 . The apparatus as defined in claim 15 , wherein the processing circuitry is to perform the plurality of partial matrix operations in a plurality of stages.
19 . The apparatus as defined in claim 15 , wherein the processing circuitry is to perform the plurality of partial matrix operations via a plurality of matrix processing units (MPUs), wherein each MPU includes processing circuitry to perform matrix arithmetic.
20 . The apparatus as defined in claim 15 , wherein the processing circuitry is to perform the plurality of partial matrix operations via performing a first processing stage via a first partition of the input matrices and performing a second processing stage via a second partition of the input matrices.
21 . The apparatus as defined in claim 15 , wherein the plurality of input matrices includes matrix data associated with one or more images.Join the waitlist — get patent alerts
Track US2022121954A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.