Hardware accelerated neural network subgraphs
Abstract
Technology related to hardware accelerated neural network subgraphs is disclosed. In one example of the disclosed technology, a method for compiling a neural network model is disclosed. The method includes identifying a subgraph of the neural network model to partition from the neural network model. An interface can be inserted between the neural network model and a partitioned version of the identified subgraph. The partitioned version can be adapted to be evaluated with a neural network accelerator. The identified subgraph can be compiled to the neural network accelerator to generate configuration information for the neural network accelerator. The neural network accelerator can be configured with the configuration information to provide an accelerated version of the subgraph.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for compiling a neural network model, comprising:
identifying a subgraph of the neural network model to partition from the neural network model; inserting an interface between the neural network model and a partitioned version of the identified subgraph, the partitioned version being adapted to be evaluated with a neural network accelerator; compiling the identified subgraph to the neural network accelerator to generate configuration information for the neural network accelerator; and configuring the neural network accelerator with the configuration information to provide an accelerated version of the subgraph.
2 . The method of claim 1 , wherein inserting the interface comprises identifying a group of edges at a boundary of the identified subgraph.
3 . The method of claim 2 , wherein inserting the interface comprises generating a data structure for passing tensor values between the neural network model and the partitioned version of the identified subgraph across the identified group of edges.
4 . The method of claim 3 , wherein generating the data structure comprises specifying an order of tensor values within the data structure, each tensor value corresponding to a different respective edge of the group of edges.
5 . The method of claim 1 , wherein compiling the identified subgraph comprises assigning training data to particular memory elements of the neural network accelerator, the training data including weights and biases corresponding to nodes of the identified subgraph.
6 . The method of claim 1 , wherein compiling the identified subgraph comprises assigning a particular region of configurable logic of the neural network accelerator to evaluate a particular neural node of the identified subgraph.
7 . The method of claim 6 , wherein compiling the identified subgraph comprises assigning training data corresponding to the particular node of the subgraph to a memory element that is locally accessible to the particular region of configurable logic of the neural network accelerator.
8 . A method for evaluating a neural network model, comprising:
using a neural network accelerator to evaluate a subgraph of the neural network model to generate output values corresponding to a first boundary of the subgraph; using a neural network server including a general-purpose central processing unit (CPU) to evaluate the neural network model to generate input values corresponding to a second boundary of the subgraph; and communicating the generated input values of the subgraph from the neural network server to the neural network accelerator using a packet comprising an identifier identifying the second boundary and the generated input values.
9 . The method of claim 8 , wherein the identifier identifying the second boundary is associated with particular memory elements of the neural network accelerator and the generated input values of the subgraph are stored in the particular memory elements in response to receiving the packet.
10 . The method of claim 9 , wherein the particular memory elements are block RAMs associated with neural node processing elements that are configured to evaluate nodes of the subgraph that are connected to the second boundary of the subgraph.
11 . The method of claim 8 , further comprising:
loading training data into particular memory elements of the neural network accelerator prior to evaluating the neural network model in an inference mode.
12 . The method of claim 11 , wherein the training data comprises weights and biases for neural nodes of the subgraph.
13 . The method of claim 8 , further comprising:
communicating the generated output values of the subgraph from the neural network accelerator to the neural network server using a packet comprising an identifier identifying the first boundary and the generated output values.
14 . A system, comprising:
a neural network server in communication with a neural network accelerator, the neural network server comprising:
at least one processor, and
a computer-readable memory storing computer-executable instructions that when executed by the at least one processor, cause the neural network server to perform a method, the instructions comprising:
instructions to compile a neural network model for execution on the system, wherein compiling the neural network model comprises partitioning a subgraph of the neural network model for execution on the neural network accelerator and generating configuration data for configuring the neural network accelerator;
instructions to, during a deployment mode, use the configuration data to configure the neural network accelerator to perform operations of the subgraph of the neural network model; and
instructions to evaluate the neural network model during an inference mode, the evaluation comprising passing tensor values between the neural network server and the neural network accelerator; and
wherein the neural network accelerator comprises:
configurable logic that is configurable using at least the generated configuration data, the configurable logic comprising a plurality of regions, a respective region configured to perform an operation of a respective node of the subgraph; and
memory comprising a plurality of memory elements, wherein a respective memory element is locally accessible by a respective region of the configurable logic.
15 . The system of claim 14 , wherein the instructions further comprise:
instructions to, during the deployment mode, load weights and a bias for a given node of the subgraph into the memory element that is locally accessible by the respective region of the configurable logic that is configured to perform operations for the given node.
16 . The system of claim 14 , wherein partitioning the subgraph of the neural network model for execution on the neural network accelerator comprises identifying input edges of the subgraph and generating a data structure for passing values from the input edges of the subgraph to neural nodes of the subgraph.
17 . The system of claim 16 , wherein the tensor values are passed between the neural network server and the neural network accelerator using a packet comprising the tensor values formatted according to the generated data structure.
18 . The system of claim 14 , wherein the tensor values are passed between the neural network server and the neural network accelerator using an application-layer packet consisting of only an identifier identifying the subgraph and the tensor values.
19 . The system of claim 14 , wherein the configurable logic of the neural network accelerator comprises support logic for broadcasting the tensor values passed to the neural network accelerator to the memory elements associated with input neural nodes of the subgraph.
20 . The system of claim 14 , wherein the configurable logic of the neural network accelerator is configured to implement a soft central processing unit (CPU) for processing at least a portion of the hardware accelerated subgraph.Join the waitlist — get patent alerts
Track US2019286972A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.