US2024028894A1PendingUtilityA1

Splitting of input data for processing in neural network processor

Assignee: APPLE INCPriority: May 4, 2018Filed: Jul 27, 2023Published: Jan 25, 2024
Est. expiryMay 4, 2038(~11.8 yrs left)· nominal 20-yr term from priority
G06N 3/08G06N 3/06G06N 3/0464G06N 7/04G06N 3/045G06N 3/063
77
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Embodiments of the present disclosure relate to splitting input data into smaller units for loading into a data buffer and neural engines in a neural processor circuit for performing neural network operations. The input data of a large size is split into slices and each slice is again split into tiles. The tile is uploaded from an external source to a data buffer inside the neural processor circuit but outside the neural engines. Each tile is again split into work units sized for storing in an input buffer circuit inside each neural engine. The input data stored in the data buffer and the input buffer circuit is reused by the neural engines to reduce re-fetching of input data. Operations of splitting the input data are performed at various components of the neural processor circuit under the management of rasterizers provided in these components.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A method, comprising:
 providing, to a plurality of neural engines, work units corresponding to a segment of input data from a data buffer circuit;   storing the work units in an input buffer circuit of at least one of the plurality of neural engines;   shifting a portion of the work units; and   performing, by at least one of the plurality of neural engines, a digital multiply and add operation on the shifted portion of the work units using a kernel.   
     
     
         3 . The method of  claim 2 , wherein the segment of the input data has a first size. 
     
     
         4 . The method of  claim 3 , further comprising:
 providing, to the plurality of neural engines, second work units corresponding to a second segment of the input data from the data buffer circuit, wherein the second segment of the input data has a second size.   
     
     
         5 . The method of  claim 2 , wherein shifting the portion of the work units comprises:
 shifting the portion of the work units based on task information, wherein the task information indicates a manner in which the input data is segmented into the work units.   
     
     
         6 . The method of  claim 5 , wherein the task information further comprises:
 a dimension of the input data.   
     
     
         7 . The method of  claim 2 , further comprising:
 generating a processed digital value based on the digital multiply and add operation on the shifted portion of the work units using the kernel; and   storing the processed digital value in an accumulator of the at least one of the plurality of neural engines.   
     
     
         8 . The method of  claim 7 , further comprising:
 performing, by at least one of the plurality of neural engines, a second digital multiply and add operation on the processed digital value.   
     
     
         9 . A neural processor circuit, comprising:
 a plurality of neural engines, each comprising a input buffer circuit;   a data buffer circuit configured to provide, to a respective input buffer circuit of at least one neural engine of the plurality of neural engines, a respective work unit of a plurality of work units, wherein each respective work unit corresponds to a portion of a segment of input data; and   a rasterizer circuit configured to instruct the respective input buffer circuit of the at least one neural engine to shift a portion of the respective work unit,
 wherein the at least one neural engine comprises a multiply-accumulator circuit configured to perform a digital multiply and add operation on the shifted portion of the respective work unit using a kernel. 
   
     
     
         10 . The neural processor circuit of  claim 9 , wherein the segment of the input data has a first size. 
     
     
         11 . The neural processor circuit of  claim 10 , wherein the data buffer circuit is further configured to:
 provide, to the respective input buffer circuit, a respective second work unit corresponding to a second segment of the input data, wherein the second segment of the input data has a second size.   
     
     
         12 . The neural processor circuit of  claim 9 , wherein the rasterizer circuit is configured to instruct the respective input buffer circuit to shift the portion of the respective work unit provided thereto based on task information, wherein the task information indicates a manner in which the input data is segmented into the respective work unit. 
     
     
         13 . The neural processor circuit of  claim 12 , wherein the task information further comprises:
 a dimension of the input data.   
     
     
         14 . The neural processor circuit of  claim 9 , wherein the multiply-accumulator circuit is further configured to:
 generate a processed digital value based on the digital multiply and add operation on the shifted portion of the respective work unit using the kernel, wherein the at least one neural engine further comprises an accumulator configured to store the processed digital value.   
     
     
         15 . The neural processor circuit of  claim 14 , wherein the multiply-accumulator circuit is further configured to:
 perform a second digital multiply and add operation on the processed digital value.   
     
     
         16 . A system, comprising:
 a data buffer circuit; and   a plurality of neural engines, each neural engine of the plurality of neural engines configured to:
 obtain a respective work unit of a plurality of work units from the data buffer circuit, wherein each respective work unit corresponds to a portion of a segment of input data; 
 store the respective work unit; 
 shift a portion of the respective work unit; and 
 perform a digital multiply and add operation on the shifted portion of the respective work unit using a kernel. 
   
     
     
         17 . The system of  claim 16 , wherein the segment of the input data has a first size. 
     
     
         18 . The system of  claim 17 , wherein each neural engine of the plurality of neural engines is further configured to obtain a respective second work unit corresponding to a second segment of the input data from the data buffer circuit, wherein the second segment of the input data has a second size. 
     
     
         19 . The system of  claim 18 , wherein, to shift the portion of the respective work unit, each neural engine of the plurality of neural engines is configured to shift the portion of the respective work unit based on task information, wherein the task information indicates a manner in which the input data is segmented into the respective work unit. 
     
     
         20 . The system of  claim 19 , wherein the task information further comprises:
 a dimension of the input data;   
     
     
         21 . The system of  claim 16 , wherein each neural engine of the plurality of neural engines is further configured to:
 generate a processed digital value based on the digital multiply and add operation on the shifted portion of the respective work unit using the kernel; and   store the processed digital value in an accumulator of the neural engine.

Join the waitlist — get patent alerts

Track US2024028894A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.