US2019065954A1PendingUtilityA1

Memory bandwidth management for deep learning applications

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Jun 25, 2015Filed: Oct 30, 2018Published: Feb 28, 2019
Est. expiryJun 25, 2035(~8.9 yrs left)· nominal 20-yr term from priority
G06N 3/044G06N 3/063G06N 3/04G06N 3/08G06N 3/0445G10L 15/16G06N 3/0499G06N 3/09G06N 3/0442G06F 9/3885
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In a data center, neural network evaluations can be included for services involving image or speech recognition by using a field programmable gate array (FPGA) or other parallel processor. The memory bandwidth limitations of providing weighted data sets from an external memory to the FPGA (or other parallel processor) can be managed by queuing up input data from the plurality of cores executing the services at the FPGA (or other parallel processor) in batches of at least two feature vectors. The at least two feature vectors can be at least two observation vectors from a same data stream or from different data streams. The FPGA (or other parallel processor) can then act on the batch of data for each loading of the weighted datasets.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of performing neural network processes, comprising:
 loading a parallel processor with a first set of weights for a neural network process from an external memory;   applying the first set of weights sequentially to at least two sets of input data such that one of the at least two sets of input data is processed after another of the at least two sets of input data while each set of input data of the at least two sets of input data is separately processed in parallel by the parallel processor;   queuing intermediates of the at least two sets of input data;   loading the parallel processor with a second set of weights for the neural network process from the external memory;   applying the second set of weights sequentially to the intermediates of the at least two sets of input data; and   repeating the queuing, loading, and applying until the neural network process is completed for the at least two sets of input data.   
     
     
         2 . The method of  claim 1 , wherein the input data comprises data for processing for speech recognition or translation services. 
     
     
         3 . The method of  claim 1 , wherein the input data comprises data for processing for computer vision applications. 
     
     
         4 . The method of  claim 1 , wherein the input data comprises data for image processing or recognition services. 
     
     
         5 . The method of  claim 1 , wherein the input data comprises data for natural language processing. 
     
     
         6 . The method of  claim 1 , wherein the input data comprises data for audio recognition. 
     
     
         7 . The method of  claim 1 , wherein the input data comprises data for bioinformatics. 
     
     
         8 . The method of  claim 1 , wherein the input data comprises data for weather prediction. 
     
     
         9 . A system comprising:
 a parallel processor having N available parallel streams of processing and a buffer for each stream having a queue depth of at least two such that one stream of parallel input data from at least two sets of parallel input data is stored in the buffer before processing on a particular weight dataset; and   storage storing weight datasets, including the particular weight dataset, for a neural network evaluation.   
     
     
         10 . The system of  claim 9 , wherein the parallel processor is a field programmable gate array (FPGA). 
     
     
         11 . The system of  claim 9 , further comprising:
 a plurality of processing cores operatively coupled to the parallel processor, the plurality of processing cores executing audio, speech, or image processing.   
     
     
         12 . The system of  claim 9 , wherein the parallel processor is operatively coupled to receive at least one observation vector for a process executing on at least one of the plurality of processing cores and communicate an evaluation output to a respective at least one of the plurality of processing cores. 
     
     
         13 . One or more storage media having instructions stored thereon, that when executed direct a computing system to:
 load a parallel processor with a first set of weights for a neural network process from an external memory, wherein the first set of weights is applied sequentially to at least two sets of input data such that one of the at least two sets of input data is processed after another of the at least two sets of input data while each set of input data of the at least two sets of input data is separately processed in parallel by the parallel processor, wherein intermediates of the at least two sets of input data are queued after the first set of weights is applied; and   load the parallel processor with a second set of weights for the neural network process from the external memory, wherein the second set of weights is applied sequentially to the intermediates of the at least two sets of input data.   
     
     
         14 . The media of  claim 13 , wherein the input data comprises data for processing for speech recognition or translation services. 
     
     
         15 . The media of  claim 13 , wherein the input data comprises data for processing for computer vision applications. 
     
     
         16 . The media of  claim 13 , wherein the input data comprises data for image processing or recognition services. 
     
     
         17 . The media of  claim 13 , wherein the input data comprises data for natural language processing. 
     
     
         18 . The media of  claim 13 , wherein the input data comprises data for audio recognition. 
     
     
         19 . The media of  claim 13 , wherein the input data comprises data for bioinformatics. 
     
     
         20 . The media of  claim 13 , wherein the input data comprises data for weather prediction.

Join the waitlist — get patent alerts

Track US2019065954A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.