US2019286972A1PendingUtilityA1

Hardware accelerated neural network subgraphs

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Mar 14, 2018Filed: May 4, 2018Published: Sep 19, 2019
Est. expiryMar 14, 2038(~11.6 yrs left)· nominal 20-yr term from priority
G06N 3/08G06N 3/063G06F 8/451G06N 3/04G06N 3/0442G06N 3/09G06N 3/0495G06N 3/0464
52
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Technology related to hardware accelerated neural network subgraphs is disclosed. In one example of the disclosed technology, a method for compiling a neural network model is disclosed. The method includes identifying a subgraph of the neural network model to partition from the neural network model. An interface can be inserted between the neural network model and a partitioned version of the identified subgraph. The partitioned version can be adapted to be evaluated with a neural network accelerator. The identified subgraph can be compiled to the neural network accelerator to generate configuration information for the neural network accelerator. The neural network accelerator can be configured with the configuration information to provide an accelerated version of the subgraph.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for compiling a neural network model, comprising:
 identifying a subgraph of the neural network model to partition from the neural network model;   inserting an interface between the neural network model and a partitioned version of the identified subgraph, the partitioned version being adapted to be evaluated with a neural network accelerator;   compiling the identified subgraph to the neural network accelerator to generate configuration information for the neural network accelerator; and   configuring the neural network accelerator with the configuration information to provide an accelerated version of the subgraph.   
     
     
         2 . The method of  claim 1 , wherein inserting the interface comprises identifying a group of edges at a boundary of the identified subgraph. 
     
     
         3 . The method of  claim 2 , wherein inserting the interface comprises generating a data structure for passing tensor values between the neural network model and the partitioned version of the identified subgraph across the identified group of edges. 
     
     
         4 . The method of  claim 3 , wherein generating the data structure comprises specifying an order of tensor values within the data structure, each tensor value corresponding to a different respective edge of the group of edges. 
     
     
         5 . The method of  claim 1 , wherein compiling the identified subgraph comprises assigning training data to particular memory elements of the neural network accelerator, the training data including weights and biases corresponding to nodes of the identified subgraph. 
     
     
         6 . The method of  claim 1 , wherein compiling the identified subgraph comprises assigning a particular region of configurable logic of the neural network accelerator to evaluate a particular neural node of the identified subgraph. 
     
     
         7 . The method of  claim 6 , wherein compiling the identified subgraph comprises assigning training data corresponding to the particular node of the subgraph to a memory element that is locally accessible to the particular region of configurable logic of the neural network accelerator. 
     
     
         8 . A method for evaluating a neural network model, comprising:
 using a neural network accelerator to evaluate a subgraph of the neural network model to generate output values corresponding to a first boundary of the subgraph;   using a neural network server including a general-purpose central processing unit (CPU) to evaluate the neural network model to generate input values corresponding to a second boundary of the subgraph; and   communicating the generated input values of the subgraph from the neural network server to the neural network accelerator using a packet comprising an identifier identifying the second boundary and the generated input values.   
     
     
         9 . The method of  claim 8 , wherein the identifier identifying the second boundary is associated with particular memory elements of the neural network accelerator and the generated input values of the subgraph are stored in the particular memory elements in response to receiving the packet. 
     
     
         10 . The method of  claim 9 , wherein the particular memory elements are block RAMs associated with neural node processing elements that are configured to evaluate nodes of the subgraph that are connected to the second boundary of the subgraph. 
     
     
         11 . The method of  claim 8 , further comprising:
 loading training data into particular memory elements of the neural network accelerator prior to evaluating the neural network model in an inference mode.   
     
     
         12 . The method of  claim 11 , wherein the training data comprises weights and biases for neural nodes of the subgraph. 
     
     
         13 . The method of  claim 8 , further comprising:
 communicating the generated output values of the subgraph from the neural network accelerator to the neural network server using a packet comprising an identifier identifying the first boundary and the generated output values.   
     
     
         14 . A system, comprising:
 a neural network server in communication with a neural network accelerator, the neural network server comprising:
 at least one processor, and 
 a computer-readable memory storing computer-executable instructions that when executed by the at least one processor, cause the neural network server to perform a method, the instructions comprising:
 instructions to compile a neural network model for execution on the system, wherein compiling the neural network model comprises partitioning a subgraph of the neural network model for execution on the neural network accelerator and generating configuration data for configuring the neural network accelerator; 
 instructions to, during a deployment mode, use the configuration data to configure the neural network accelerator to perform operations of the subgraph of the neural network model; and 
 instructions to evaluate the neural network model during an inference mode, the evaluation comprising passing tensor values between the neural network server and the neural network accelerator; and 
 
   wherein the neural network accelerator comprises:
 configurable logic that is configurable using at least the generated configuration data, the configurable logic comprising a plurality of regions, a respective region configured to perform an operation of a respective node of the subgraph; and 
 memory comprising a plurality of memory elements, wherein a respective memory element is locally accessible by a respective region of the configurable logic. 
   
     
     
         15 . The system of  claim 14 , wherein the instructions further comprise:
 instructions to, during the deployment mode, load weights and a bias for a given node of the subgraph into the memory element that is locally accessible by the respective region of the configurable logic that is configured to perform operations for the given node.   
     
     
         16 . The system of  claim 14 , wherein partitioning the subgraph of the neural network model for execution on the neural network accelerator comprises identifying input edges of the subgraph and generating a data structure for passing values from the input edges of the subgraph to neural nodes of the subgraph. 
     
     
         17 . The system of  claim 16 , wherein the tensor values are passed between the neural network server and the neural network accelerator using a packet comprising the tensor values formatted according to the generated data structure. 
     
     
         18 . The system of  claim 14 , wherein the tensor values are passed between the neural network server and the neural network accelerator using an application-layer packet consisting of only an identifier identifying the subgraph and the tensor values. 
     
     
         19 . The system of  claim 14 , wherein the configurable logic of the neural network accelerator comprises support logic for broadcasting the tensor values passed to the neural network accelerator to the memory elements associated with input neural nodes of the subgraph. 
     
     
         20 . The system of  claim 14 , wherein the configurable logic of the neural network accelerator is configured to implement a soft central processing unit (CPU) for processing at least a portion of the hardware accelerated subgraph.

Join the waitlist — get patent alerts

Track US2019286972A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.