US2025045562A1PendingUtilityA1
Method and apparatus with neural network execution
Est. expiryMar 7, 2043(~16.6 yrs left)· nominal 20-yr term from priority
G06N 3/063G06F 2212/221G06F 3/0659G06N 3/082G06N 3/0464G06N 3/0442
56
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A processor-implemented method includes: determining an operation sequence of a neural network based on dependency information of a tile constituting a feature map of the neural network and layer information of the neural network; and generating a first command for controlling a feature map memory and a second command for controlling an operator based on the operation sequence, wherein the first command comprises information on a tile input to each of a plurality of memory queues constituting the feature map memory.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A processor-implemented method comprising:
determining an operation sequence of a neural network based on dependency information of a tile constituting a feature map of the neural network and layer information of the neural network; and generating a first command for controlling a feature map memory and a second command for controlling an operator based on the operation sequence, wherein the first command comprises information on a tile input to each of a plurality of memory queues constituting the feature map memory.
2 . The method of claim 1 , wherein the dependency information of the tile comprises information on another tile that is used to perform an operation on a predetermined tile and comprises an overlap region.
3 . The method of claim 1 , wherein the second command comprises information on a tile comprised in a memory queue that is a target of an operation among the plurality of memory queues for the operator.
4 . The method of claim 1 , wherein the memory queues are separated and constituted based on a layer of the neural network.
5 . The method of claim 1 , wherein each of the plurality of memory queues is constituted in a plurality of memory banks stored in a unit of tiles.
6 . The method of claim 1 , wherein a tile constituting the feature map of the neural network is generated by dividing the feature map in a single direction.
7 . The method of claim 5 , wherein, in response to a bit width in a first direction of the feature map being greater than a bit width of the memory bank, a bit width of the tile is determined based on the bit width of the memory bank.
8 . The method of claim 1 , wherein the feature map memory comprises static random access memory (SRAM).
9 . A processor-implemented method comprising:
obtaining information of a tile constituting a feature map of a neural network; determining an operation sequence of the neural network based on the information of the tile; determining a first tile to be stored in a feature map memory based on the operation sequence; storing the first tile and a second tile that is used to perform an operation on the first tile and comprises an overlap region in a first memory queue of the feature map memory; performing an operation on the first tile based on the first tile and the second tile; and storing an operation result in a second memory queue of the feature map memory.
10 . The method of claim 9 , wherein the determining of the operation sequence comprises determining the operation sequence based on dependency information of the tile.
11 . The method of claim 9 , wherein the information of the tile is determined based on a shared boundary area of a tile constituting the feature map.
12 . A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform the method of claim 1 .
13 . An apparatus comprising:
a memory; an operator; and a scheduler configured to:
determine an operation sequence of a neural network based on dependency information of a tile constituting a feature map of the neural network and layer information of the neural network; and
generate a first command for controlling the memory and a second command for controlling the operator based on the operation sequence,
wherein the first command comprises information on a tile input to each of a plurality of memory queues constituting the memory.
14 . The apparatus of claim 13 , wherein the dependency information of the tile comprises information on another tile that is used to perform an operation on a predetermined tile and comprises an overlap region.
15 . The apparatus of claim 13 , wherein the second command comprises information on a first tile comprised in a first memory queue that is a target of an operation among the plurality of memory queues for the operator.
16 . The apparatus of claim 15 , wherein
the operator is configured to generate a second tile by processing the first tile based on the second command, and the memory is configured to store the second tile in a second memory queue.
17 . The apparatus of claim 13 , wherein the memory queues are separated and constituted based on a layer of the neural network.
18 . The apparatus of claim 13 , wherein each of the plurality of memory queues is constituted in a plurality of memory banks stored in a unit of tiles.
19 . The apparatus of claim 13 , wherein a tile constituting the feature map of the neural network is generated by dividing the feature map in a single direction.
20 . The apparatus of claim 18 , wherein, in response to a bit width in a first direction of the feature map being greater than a bit width of the memory bank, a bit width of the tile is determined based on the bit width of the memory bank.Join the waitlist — get patent alerts
Track US2025045562A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.