US2025045562A1PendingUtilityA1

Method and apparatus with neural network execution

Assignee: SAMSUNG ELECTRONICS CO LTDPriority: Mar 7, 2023Filed: Aug 4, 2023Published: Feb 6, 2025
Est. expiryMar 7, 2043(~16.6 yrs left)· nominal 20-yr term from priority
G06N 3/063G06F 2212/221G06F 3/0659G06N 3/082G06N 3/0464G06N 3/0442
56
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A processor-implemented method includes: determining an operation sequence of a neural network based on dependency information of a tile constituting a feature map of the neural network and layer information of the neural network; and generating a first command for controlling a feature map memory and a second command for controlling an operator based on the operation sequence, wherein the first command comprises information on a tile input to each of a plurality of memory queues constituting the feature map memory.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A processor-implemented method comprising:
 determining an operation sequence of a neural network based on dependency information of a tile constituting a feature map of the neural network and layer information of the neural network; and   generating a first command for controlling a feature map memory and a second command for controlling an operator based on the operation sequence,   wherein the first command comprises information on a tile input to each of a plurality of memory queues constituting the feature map memory.   
     
     
         2 . The method of  claim 1 , wherein the dependency information of the tile comprises information on another tile that is used to perform an operation on a predetermined tile and comprises an overlap region. 
     
     
         3 . The method of  claim 1 , wherein the second command comprises information on a tile comprised in a memory queue that is a target of an operation among the plurality of memory queues for the operator. 
     
     
         4 . The method of  claim 1 , wherein the memory queues are separated and constituted based on a layer of the neural network. 
     
     
         5 . The method of  claim 1 , wherein each of the plurality of memory queues is constituted in a plurality of memory banks stored in a unit of tiles. 
     
     
         6 . The method of  claim 1 , wherein a tile constituting the feature map of the neural network is generated by dividing the feature map in a single direction. 
     
     
         7 . The method of  claim 5 , wherein, in response to a bit width in a first direction of the feature map being greater than a bit width of the memory bank, a bit width of the tile is determined based on the bit width of the memory bank. 
     
     
         8 . The method of  claim 1 , wherein the feature map memory comprises static random access memory (SRAM). 
     
     
         9 . A processor-implemented method comprising:
 obtaining information of a tile constituting a feature map of a neural network;   determining an operation sequence of the neural network based on the information of the tile;   determining a first tile to be stored in a feature map memory based on the operation sequence;   storing the first tile and a second tile that is used to perform an operation on the first tile and comprises an overlap region in a first memory queue of the feature map memory;   performing an operation on the first tile based on the first tile and the second tile; and   storing an operation result in a second memory queue of the feature map memory.   
     
     
         10 . The method of  claim 9 , wherein the determining of the operation sequence comprises determining the operation sequence based on dependency information of the tile. 
     
     
         11 . The method of  claim 9 , wherein the information of the tile is determined based on a shared boundary area of a tile constituting the feature map. 
     
     
         12 . A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform the method of  claim 1 . 
     
     
         13 . An apparatus comprising:
 a memory;   an operator; and   a scheduler configured to:
 determine an operation sequence of a neural network based on dependency information of a tile constituting a feature map of the neural network and layer information of the neural network; and 
 generate a first command for controlling the memory and a second command for controlling the operator based on the operation sequence, 
   wherein the first command comprises information on a tile input to each of a plurality of memory queues constituting the memory.   
     
     
         14 . The apparatus of  claim 13 , wherein the dependency information of the tile comprises information on another tile that is used to perform an operation on a predetermined tile and comprises an overlap region. 
     
     
         15 . The apparatus of  claim 13 , wherein the second command comprises information on a first tile comprised in a first memory queue that is a target of an operation among the plurality of memory queues for the operator. 
     
     
         16 . The apparatus of  claim 15 , wherein
 the operator is configured to generate a second tile by processing the first tile based on the second command, and   the memory is configured to store the second tile in a second memory queue.   
     
     
         17 . The apparatus of  claim 13 , wherein the memory queues are separated and constituted based on a layer of the neural network. 
     
     
         18 . The apparatus of  claim 13 , wherein each of the plurality of memory queues is constituted in a plurality of memory banks stored in a unit of tiles. 
     
     
         19 . The apparatus of  claim 13 , wherein a tile constituting the feature map of the neural network is generated by dividing the feature map in a single direction. 
     
     
         20 . The apparatus of  claim 18 , wherein, in response to a bit width in a first direction of the feature map being greater than a bit width of the memory bank, a bit width of the tile is determined based on the bit width of the memory bank.

Join the waitlist — get patent alerts

Track US2025045562A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.