Configurable pooling processing unit for neural network accelerator
Abstract
A hardware implementation of a configurable pooling processing unit is configured to receive an input tensor comprising at least one channel, each channel of the at least one channel comprising a plurality of tensels; receive control information identifying one operation of a plurality of selectable operations to be performed on the input tensor, the plurality of selectable operations comprising a depth-wise convolution operation and one or more pooling operations; perform the identified operation on the input tensor to generate an output tensor by performing one or more operations on blocks of tensels of each channel of the at least one channel of the input tensor; and output the output tensor.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A hardware accelerator to implement a configurable pooling processing unit, the hardware accelerator configured to:
receive an input tensor comprising at least one channel, each channel of the at least one channel comprising a plurality of tensels; receive control information identifying one operation of a plurality of selectable operations to be performed on the input tensor, the plurality of selectable operations comprising a depth-wise convolution operation and one or more pooling operations; perform, using a same set of hardware components of the hardware accelerator regardless of the identified operation, the identified operation on the input tensor to generate an output tensor by performing one or more operations on blocks of tensels of each channel of the at least one channel of the input tensor; and output the output tensor.
2 . The hardware accelerator of claim 1 , comprising one or more pool engines, each pool engine configurable to receive tensels of a channel of the input tensor and generate one of a plurality of different types of channel outputs, the plurality of different types of channel outputs comprising a depth-wise convolution channel output and one of one or more selectable pooling operation channel outputs.
3 . The hardware accelerator of claim 2 , wherein the one or more pooling operations comprises an average pooling operation.
4 . The hardware accelerator of claim 3 , wherein each of the one or more pool engines comprises:
a reduction engine configurable to generate, for a block of tensels of a channel of the input tensor, one of a plurality of types of block outputs, the plurality of types of block outputs comprising a sum of tensels in the block and a weighted sum of tensels in the block; and a division engine configurable to selectively perform a division operation on the block output generated by the reduction engine; wherein when the control information identifies that an average pooling operation is to be performed on the input tensor, the reduction engine is configured to generate a sum of tensels in the block and the division engine is enabled to divide the block output generated by the reduction engine by a number of tensels in the block; and wherein when the control information identifies that a depth-wise convolution operation is to be performed on the input tensor, the reduction engine is configured to generate a weighted sum for the block and the division engine is disabled.
5 . The hardware accelerator of claim 4 , wherein each block of tensels comprises one or more rows of tensels and one or more columns of tensels, and the reduction engine is configured to generate a block output by generating column outputs and generating the block output from one or more column outputs.
6 . The hardware accelerator of claim 5 , wherein:
when the control information identifies that an average pooling operation is to be performed on the input tensor, the reduction engine is configured to generate a sum for each column of a block of tensels, and generate the sum for the block of tensels by summing appropriate column sums, and when the control information identifies that a depth-wise convolution operation is to be performed on the input tensor, the reduction engine is configured to generate a weighted sum for each column of a block of tensels, and generate the weighted sum for the block by summing appropriate column weighted sums.
7 . The hardware accelerator of claim 4 , wherein the reduction engine comprises:
a vertical pool engine configurable to receive a column of tensels and generate one of a plurality of types of column outputs for that column; a collector storage unit configured to temporarily store the column outputs generated by the vertical pool engine; and a horizontal pool engine configured to generate a block output from appropriate column outputs stored in the collector storage unit.
8 . The hardware accelerator of claim 7 , wherein:
when the control information identifies that an average pooling operation is to be performed on the input tensor, the vertical pool engine is configured to receive a column of tensels in a block and generate a sum of the received tensels; and when the control information identifies that a depth-wise convolution operation is to be performed on the input tensor, the vertical pool engine is configured to receive a column of tensels in a block, and generate a plurality of weighted sums for the received tensels, each weighted sum based on a different set of weights.
9 . The hardware accelerator of claim 7 , wherein the vertical pool engine comprises:
a plurality of multiplication units, each multiplication unit configurable to receive a set of multiplication input elements and multiply each of the received multiplication input elements with a corresponding weight to generate a multiplication output; and a plurality of summation units, each summation unit configurable to receive a set of summation input elements and generate a sum of the received summation input elements to generate a summation output; wherein when the control information identifies that an average pooling operation is to be performed on the input tensor, one of the plurality of summation units is configured to receive a set tensels in a column and generate the sum of the set of tensels; and wherein when the control information identifies that a depth-wise convolution operation is to be performed on the input tensor, at least two of the plurality of multiplication units are configured to receive a same set of tensels in a column and generate multiplication outputs based on a different set of weights, and at least two of the plurality of summation units are configured to generate a sum of the multiplication outputs for one of the at least two multiplication units.
10 . The hardware accelerator of claim 9 , wherein each set of weights corresponds to a column of a filter to be applied to a channel of the input tensor.
11 . The hardware accelerator of claim 7 , wherein the collector storage unit is a register, and a set of pointers identify the appropriate column outputs in the register to generate a block output.
12 . The hardware accelerator of claim 4 , wherein each pool engine further comprises a post calculation engine configurable to reformat an output of the reduction engine or the division engine.
13 . The hardware accelerator of claim 1 , comprising a parameter storage unit, and when the control information identifies that a depth-wise convolution operation is to be performed on the input tensor, the hardware accelerator is configured to fetch parameters for performing the depth-wise convolution operation and store the fetched parameters in the parameter storage unit, the parameters for performing the depth-wise convolution operation comprising a set of parameters for each channel of the at least one channel of the input tensor, the set of parameters for a channel comprising a set of weights.
14 . The hardware accelerator of claim 13 , wherein the set of parameters for a channel further comprises a bias value.
15 . The hardware accelerator of claim 13 , wherein, when the weights for a channel are in an affine fixed point number format, the set of parameters for a channel further comprise a weight zero point, and the hardware accelerator is configured to remove the weight zero point from each weight associated with that channel prior to performing the depth-wise convolution operation.
16 . The hardware accelerator of claim 1 , wherein the hardware accelerator is embodied on an integrated circuit.
17 . A neural network accelerator comprising the hardware accelerator as set forth in claim 1 .
18 . The neural network accelerator of claim 17 , further comprising a convolution processing unit configurable to perform one of a plurality of different convolution operations.
19 . A non-transitory computer readable storage medium having stored thereon a computer readable dataset description of the hardware accelerator as set forth in claim 1 that, when processed in an integrated circuit manufacturing system, causes the integrated circuit manufacturing system to manufacture an integrated circuit embodying the hardware accelerator.
20 . A method of processing, at a hardware accelerator configured to implement a configurable pooling processing unit, an input tensor comprising at least one channel, each channel of the at least one channel comprising a plurality of tensels, the method comprising:
receiving control information identifying one operation of a plurality of selectable operations to be performed on the input tensor, the plurality of selectable operations comprising a depth-wise convolution operation and one or more pooling operations; performing, using a same set of hardware components of the hardware accelerator regardless of the identified operation, the identified operation on the input tensor to generate an output tensor by performing one or more operations on blocks of tensels of each channel of the at least one channel of the input tensor; and outputting the output tensor.Join the waitlist — get patent alerts
Track US2023259578A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.