US2024412781A1PendingUtilityA1

Compute-in-memory processor supporting both general-purpose cpu and deep learning

Assignee: UNIV NORTHWESTERNPriority: Jun 7, 2023Filed: Jun 7, 2024Published: Dec 12, 2024
Est. expiryJun 7, 2043(~16.8 yrs left)· nominal 20-yr term from priority
Inventors:Jie GuYuhao Ju
G11C 11/54G11C 11/419G06F 7/501
42
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In certain aspects, a compute-in-memory processor includes central computing units configured to operate in a central processing unit mode and a deep neural network mode. A data activation memory and a data cache output memory are in communication with the compute-in-memory processor. In the deep neural network mode, the data activation memory is configured as input memory and the data cache output memory is configured as output memory. In the central processing unit mode, the data activation memory is configured as a first data cache and the data cache output memory is configured as a register file and a second data cache.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A compute-in-memory processor, comprising:
 central computing units configured to operate in a central processing unit mode and a deep neural network mode;   a data activation memory in communication with the central computing units; and   a data cache output memory in communication with the central computing units, wherein, in the deep neural network mode, the data activation memory is configured as input memory and the data cache output memory is configured as output memory, and wherein, in the central processing unit mode, the data activation memory is configured as a first data cache and the data cache output memory is configured as a register file and a second data cache.   
     
     
         2 . The compute-in-memory processor of  claim 1 , wherein the data activation memory comprises a bitcell array configured to support, in the deep neural network mode, SRAM function and 1b multiplication. 
     
     
         3 . The compute-in-memory processor of  claim 2 , wherein the bitcell array is a 32 bit 9 transistor bitcell array. 
     
     
         4 . The compute-in-memory processor of  claim 3 , wherein the bitcell array comprises a 3 transistor NAND gate appended to a 6 transistor SRAM. 
     
     
         5 . The compute-in-memory processor of  claim 1 , wherein the data cache output memory comprises a bitcell array, wherein the bitcell array comprises 2 bitlines configured to perform 2 read operations and 1 write operation within one clock cycle. 
     
     
         6 . The compute-in-memory processor of  claim 5 , wherein the bitcell array is an 8 transistor bitcell array. 
     
     
         7 . The compute-in-memory processor of  claim 1 , wherein the central computing units comprise a plurality of adder trees configured to perform 8 bit MAC based on 1b results from the data activation memory. 
     
     
         8 . The compute-in-memory processor of  claim 7 , wherein the plurality of adder trees comprise four adder trees. 
     
     
         9 . The compute-in-memory processor of  claim 1 , further comprising a customizable instruction set architecture configured to support integer vector CPU function in the central processing unit mode. 
     
     
         10 . The compute-in-memory processor of  claim 9 , wherein the customizable instruction set architecture is a customizable 32b instruction set architecture. 
     
     
         11 . A computer-implemented method, comprising:
 operating central computing units in a central processing unit mode; wherein, in the central processing unit mode, a data activation memory in communication with the central computing units is configured as a first data cache and a data cache output memory, in communication with the central computing units, is configured as a register file and a second data cache; and   selectively operating the central computing units from the central processing unit mode to a deep neural network mode, wherein, in the deep neural network mode, the data activation memory is configured as input memory and the data cache output memory is configured as output memory.   
     
     
         12 . The computer-implemented method of  claim 11 , wherein the data activation memory comprises a bitcell array configured to support, in the deep neural network mode, SRAM function and 1b multiplication. 
     
     
         13 . The computer-implemented method of  claim 12 , wherein the bitcell array is a 32 bit 9 transistor bitcell array. 
     
     
         14 . The computer-implemented method of  claim 13 , wherein the bitcell array comprises a 3 transistor NAND gate appended to a 6 transistor SRAM. 
     
     
         15 . The computer-implemented method of  claim 11 , wherein the data cache output memory comprises a bitcell array, wherein the bitcell array comprises 2 bitlines configured to perform 2 read operations and 1 write operation within one clock cycle. 
     
     
         16 . The computer-implemented method of  claim 11 , wherein the central computing units comprise a plurality of adder trees configured to perform 8 bit MAC based on 1b results from the data activation memory. 
     
     
         17 . The computer-implemented method of  claim 11 , wherein, in the central processing unit mode, the central computing unit is configured to support integer vector CPU function via a customizable instruction set architecture. 
     
     
         18 . The computer-implemented method of  claim 17 , wherein the customizable instruction set architecture is a customizable 32b instruction set architecture. 
     
     
         19 . A compute-in-memory processor, comprising:
 central computing units configured to operate in a central processing unit mode and a deep neural network mode;   a first level cache in communication with the central computing units; and   a second level cache in communication with the central computing units, wherein the second level cache comprises a plurality of SRAM banks, wherein, in the central processing unit mode, a second level compute-in-memory in communication with the second level cache is configured as an instruction controlled compute-in-memory core, and wherein, in the deep neural network mode, the second level compute-in-memory is configured as a weight memory for compute-in-memory MAC operations.   
     
     
         20 . The compute-in-memory processor of  claim 19 , wherein, in the deep neural network mode, a near-array vector ALU in communication with the second level cache is configured to a plurality of adder trees for partial sum accumulation.

Join the waitlist — get patent alerts

Track US2024412781A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.