US2021064379A1PendingUtilityA1
Refactoring MAC Computations for Reduced Programming Steps
Est. expiryAug 29, 2039(~13.1 yrs left)· nominal 20-yr term from priority
G06F 17/16G06F 7/5443G06N 3/06G06F 9/3893G06F 7/4876G06F 2207/4814G06F 9/30014
58
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A method and architecture for performing multiply-accumulate operations in a neural network is disclosed. The architecture includes a crossbar having a plurality of non-volatile memory elements. A plurality of input activations is applied to the crossbar, which are then summed by binary weight encoding a plurality of the non-volatile memory elements to connect the input activations to weight values. At least one of the plurality of non-volatile memory elements is then precision programmed.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of performing multiply-accumulate acceleration in a neural network, comprising:
generating, in a summing array having a plurality of non-volatile memory elements arranged in columns, a summed signal by the columns of non-volatile memory elements in the summing array, each non-volatile memory element in the summing array being programmed to either a high or low resistance state; and inputting the summed signal from the summing array to a multiplying array having a plurality of non-volatile memory elements, each non-volatile memory element in the multiplying array being precisely programmed to a conductance level proportional to a weight in the neural network.
2 . The method of claim 1 , where the summing array and multiplying array is an M×N crossbar having K-bit weights, where M×N×2 K elements are programmed to either the high or low resistance state and 2 K ×N elements are precisely programmed to the conductance level proportional to the weight in the neural network.
3 . The method of claim 2 , where a plurality of input activations are conditionally summed depending upon specific weight values.
4 . The method of claim 2 , where a plurality of input activations are significantly greater than a plurality of weight values.
5 . The method of claim 2 , where summing array comprises a plurality of high and low resistance levels.
6 . The method of claim 5 , further comprising the M×N×2 K elements programmed to either a resistance “off” state or a resistance “on” state for 2 K levels.
7 . The method of claim 6 , where 2 K ×N<<M×N non-volatile memory cells are fine-tuned.
8 . The method of claim 2 , further comprising scaling an output in a multiplier/scaling module.
9 . An architecture for performing multiply-accumulate operations in a neural network, comprising:
a summing array having a plurality of non-volatile memory elements arranged in columns, the summing array generating a summed signal by the columns of non-volatile memory elements in the summing array, each non-volatile memory element in the summing array being programmed to either a high or low resistance state; and a multiplying array having a plurality of non-volatile memory elements that receive a summed signal from the summing array, each non-volatile memory element in the multiplying array being precisely programmed to a conductance level proportional to a weight in the neural network.
10 . The architecture of claim 9 , where the summing array and multiplying array is an M×N crossbar having K-bit weights, where M×N×2 K elements are programmed to either the high or low resistance state and 2 K ×N elements are precisely programmed to the conductance level proportional to the weight in the neural network.
11 . The architecture of claim 10 , where a plurality of input activations is conditionally summed depending upon specific weight values.
12 . The architecture of claim 10 , where a plurality of input activations is significantly greater than a plurality of weight values.
13 . The architecture of claim 10 , further comprising a plurality of resistors and where the summing array comprises a plurality of high and low resistance levels.
14 . The architecture of claim 13 , further comprising the M×N×2 K elements programmed to either a resistance “off” state or a resistance “on” state for 2 K levels.
15 . The architecture of claim 14 , where 2 K ×N<<M×N non-volatile memory cells are fine-tuned.
16 . The architecture of claim 10 , further comprising a multiplier/scaling module for scaling an output.
17 . An architecture for performing multiply-accumulate operations in a neural network, comprising:
a crossbar including a plurality of crossbar nodes arranged in an array of rows and columns, each crossbar node being programmable to a first resistance level or a second resistance level, the crossbar being configured to sum a plurality of analog input activation signals over each column of crossbar nodes and output a plurality of summed activation signals; and a multiplier, coupled to the crossbar, including a plurality of multiplier nodes, each multiplier node being programmable to a resistance level proportional to one of a plurality of neural network weights, the multiplier being configured to sum the plurality of summed activation signals over the multiplier nodes and output an analog output activation signal.
18 . The architecture of claim 17 , where each crossbar node includes one or more non-volatile elements (NVMs), and each multiplier node includes a plurality of NVMs.
19 . The architecture of claim 18 , where the crossbar includes M rows, N columns, K-bit weights and M×N×2 K programmable NVMs, and the multiplier includes N multiplier nodes and N×2 K programmable NVMs.
20 . The architecture of claim 17 , further comprising:
a plurality of digital-to-analog converters (DACs) coupled to the crossbar, each DAC being configured to receive a plurality of digital input activation signals and output the plurality of analog input activation signals.Join the waitlist — get patent alerts
Track US2021064379A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.