Splitting of input data for processing in neural network processor
Abstract
Embodiments of the present disclosure relate to splitting input data into smaller units for loading into a data buffer and neural engines in a neural processor circuit for performing neural network operations. The input data of a large size is split into slices and each slice is again split into tiles. The tile is uploaded from an external source to a data buffer inside the neural processor circuit but outside the neural engines. Each tile is again split into work units sized for storing in an input buffer circuit inside each neural engine. The input data stored in the data buffer and the input buffer circuit is reused by the neural engines to reduce re-fetching of input data. Operations of splitting the input data are performed at various components of the neural processor circuit under the management of rasterizers provided in these components.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A method, comprising:
providing, to a plurality of neural engines, work units corresponding to a segment of input data from a data buffer circuit; storing the work units in an input buffer circuit of at least one of the plurality of neural engines; shifting a portion of the work units; and performing, by at least one of the plurality of neural engines, a digital multiply and add operation on the shifted portion of the work units using a kernel.
3 . The method of claim 2 , wherein the segment of the input data has a first size.
4 . The method of claim 3 , further comprising:
providing, to the plurality of neural engines, second work units corresponding to a second segment of the input data from the data buffer circuit, wherein the second segment of the input data has a second size.
5 . The method of claim 2 , wherein shifting the portion of the work units comprises:
shifting the portion of the work units based on task information, wherein the task information indicates a manner in which the input data is segmented into the work units.
6 . The method of claim 5 , wherein the task information further comprises:
a dimension of the input data.
7 . The method of claim 2 , further comprising:
generating a processed digital value based on the digital multiply and add operation on the shifted portion of the work units using the kernel; and storing the processed digital value in an accumulator of the at least one of the plurality of neural engines.
8 . The method of claim 7 , further comprising:
performing, by at least one of the plurality of neural engines, a second digital multiply and add operation on the processed digital value.
9 . A neural processor circuit, comprising:
a plurality of neural engines, each comprising a input buffer circuit; a data buffer circuit configured to provide, to a respective input buffer circuit of at least one neural engine of the plurality of neural engines, a respective work unit of a plurality of work units, wherein each respective work unit corresponds to a portion of a segment of input data; and a rasterizer circuit configured to instruct the respective input buffer circuit of the at least one neural engine to shift a portion of the respective work unit,
wherein the at least one neural engine comprises a multiply-accumulator circuit configured to perform a digital multiply and add operation on the shifted portion of the respective work unit using a kernel.
10 . The neural processor circuit of claim 9 , wherein the segment of the input data has a first size.
11 . The neural processor circuit of claim 10 , wherein the data buffer circuit is further configured to:
provide, to the respective input buffer circuit, a respective second work unit corresponding to a second segment of the input data, wherein the second segment of the input data has a second size.
12 . The neural processor circuit of claim 9 , wherein the rasterizer circuit is configured to instruct the respective input buffer circuit to shift the portion of the respective work unit provided thereto based on task information, wherein the task information indicates a manner in which the input data is segmented into the respective work unit.
13 . The neural processor circuit of claim 12 , wherein the task information further comprises:
a dimension of the input data.
14 . The neural processor circuit of claim 9 , wherein the multiply-accumulator circuit is further configured to:
generate a processed digital value based on the digital multiply and add operation on the shifted portion of the respective work unit using the kernel, wherein the at least one neural engine further comprises an accumulator configured to store the processed digital value.
15 . The neural processor circuit of claim 14 , wherein the multiply-accumulator circuit is further configured to:
perform a second digital multiply and add operation on the processed digital value.
16 . A system, comprising:
a data buffer circuit; and a plurality of neural engines, each neural engine of the plurality of neural engines configured to:
obtain a respective work unit of a plurality of work units from the data buffer circuit, wherein each respective work unit corresponds to a portion of a segment of input data;
store the respective work unit;
shift a portion of the respective work unit; and
perform a digital multiply and add operation on the shifted portion of the respective work unit using a kernel.
17 . The system of claim 16 , wherein the segment of the input data has a first size.
18 . The system of claim 17 , wherein each neural engine of the plurality of neural engines is further configured to obtain a respective second work unit corresponding to a second segment of the input data from the data buffer circuit, wherein the second segment of the input data has a second size.
19 . The system of claim 18 , wherein, to shift the portion of the respective work unit, each neural engine of the plurality of neural engines is configured to shift the portion of the respective work unit based on task information, wherein the task information indicates a manner in which the input data is segmented into the respective work unit.
20 . The system of claim 19 , wherein the task information further comprises:
a dimension of the input data;
21 . The system of claim 16 , wherein each neural engine of the plurality of neural engines is further configured to:
generate a processed digital value based on the digital multiply and add operation on the shifted portion of the respective work unit using the kernel; and store the processed digital value in an accumulator of the neural engine.Join the waitlist — get patent alerts
Track US2024028894A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.