US2021064379A1PendingUtilityA1

Refactoring MAC Computations for Reduced Programming Steps

Assignee: ADVANCED RISC MACH LTDPriority: Aug 29, 2019Filed: Aug 29, 2019Published: Mar 4, 2021
Est. expiryAug 29, 2039(~13.1 yrs left)· nominal 20-yr term from priority
G06F 17/16G06F 7/5443G06N 3/06G06F 9/3893G06F 7/4876G06F 2207/4814G06F 9/30014
58
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and architecture for performing multiply-accumulate operations in a neural network is disclosed. The architecture includes a crossbar having a plurality of non-volatile memory elements. A plurality of input activations is applied to the crossbar, which are then summed by binary weight encoding a plurality of the non-volatile memory elements to connect the input activations to weight values. At least one of the plurality of non-volatile memory elements is then precision programmed.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of performing multiply-accumulate acceleration in a neural network, comprising:
 generating, in a summing array having a plurality of non-volatile memory elements arranged in columns, a summed signal by the columns of non-volatile memory elements in the summing array, each non-volatile memory element in the summing array being programmed to either a high or low resistance state; and   inputting the summed signal from the summing array to a multiplying array having a plurality of non-volatile memory elements, each non-volatile memory element in the multiplying array being precisely programmed to a conductance level proportional to a weight in the neural network.   
     
     
         2 . The method of  claim 1 , where the summing array and multiplying array is an M×N crossbar having K-bit weights, where M×N×2 K  elements are programmed to either the high or low resistance state and 2 K ×N elements are precisely programmed to the conductance level proportional to the weight in the neural network. 
     
     
         3 . The method of  claim 2 , where a plurality of input activations are conditionally summed depending upon specific weight values. 
     
     
         4 . The method of  claim 2 , where a plurality of input activations are significantly greater than a plurality of weight values. 
     
     
         5 . The method of  claim 2 , where summing array comprises a plurality of high and low resistance levels. 
     
     
         6 . The method of  claim 5 , further comprising the M×N×2 K  elements programmed to either a resistance “off” state or a resistance “on” state for 2 K  levels. 
     
     
         7 . The method of  claim 6 , where 2 K ×N<<M×N non-volatile memory cells are fine-tuned. 
     
     
         8 . The method of  claim 2 , further comprising scaling an output in a multiplier/scaling module. 
     
     
         9 . An architecture for performing multiply-accumulate operations in a neural network, comprising:
 a summing array having a plurality of non-volatile memory elements arranged in columns, the summing array generating a summed signal by the columns of non-volatile memory elements in the summing array, each non-volatile memory element in the summing array being programmed to either a high or low resistance state; and   a multiplying array having a plurality of non-volatile memory elements that receive a summed signal from the summing array, each non-volatile memory element in the multiplying array being precisely programmed to a conductance level proportional to a weight in the neural network.   
     
     
         10 . The architecture of  claim 9 , where the summing array and multiplying array is an M×N crossbar having K-bit weights, where M×N×2 K  elements are programmed to either the high or low resistance state and 2 K ×N elements are precisely programmed to the conductance level proportional to the weight in the neural network. 
     
     
         11 . The architecture of  claim 10 , where a plurality of input activations is conditionally summed depending upon specific weight values. 
     
     
         12 . The architecture of  claim 10 , where a plurality of input activations is significantly greater than a plurality of weight values. 
     
     
         13 . The architecture of  claim 10 , further comprising a plurality of resistors and where the summing array comprises a plurality of high and low resistance levels. 
     
     
         14 . The architecture of  claim 13 , further comprising the M×N×2 K  elements programmed to either a resistance “off” state or a resistance “on” state for 2 K  levels. 
     
     
         15 . The architecture of  claim 14 , where 2 K ×N<<M×N non-volatile memory cells are fine-tuned. 
     
     
         16 . The architecture of  claim 10 , further comprising a multiplier/scaling module for scaling an output. 
     
     
         17 . An architecture for performing multiply-accumulate operations in a neural network, comprising:
 a crossbar including a plurality of crossbar nodes arranged in an array of rows and columns, each crossbar node being programmable to a first resistance level or a second resistance level, the crossbar being configured to sum a plurality of analog input activation signals over each column of crossbar nodes and output a plurality of summed activation signals; and   a multiplier, coupled to the crossbar, including a plurality of multiplier nodes, each multiplier node being programmable to a resistance level proportional to one of a plurality of neural network weights, the multiplier being configured to sum the plurality of summed activation signals over the multiplier nodes and output an analog output activation signal.   
     
     
         18 . The architecture of  claim 17 , where each crossbar node includes one or more non-volatile elements (NVMs), and each multiplier node includes a plurality of NVMs. 
     
     
         19 . The architecture of  claim 18 , where the crossbar includes M rows, N columns, K-bit weights and M×N×2 K  programmable NVMs, and the multiplier includes N multiplier nodes and N×2 K  programmable NVMs. 
     
     
         20 . The architecture of  claim 17 , further comprising:
 a plurality of digital-to-analog converters (DACs) coupled to the crossbar, each DAC being configured to receive a plurality of digital input activation signals and output the plurality of analog input activation signals.

Join the waitlist — get patent alerts

Track US2021064379A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.