US2023251861A1PendingUtilityA1

Accelerating linear algebra kernels for any processor architecture

Assignee: NVIDIA CORPPriority: Mar 9, 2018Filed: Apr 18, 2023Published: Aug 10, 2023
Est. expiryMar 9, 2038(~11.6 yrs left)· nominal 20-yr term from priority
G06F 9/3001G06F 9/30065G06T 1/20G06F 8/447
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for obtaining a set of instructions for executing a computer program and generating executable code for the computer program based, at least in part, on scheduling operations associated with the executable code according to a polyhedral representation of a directed acyclic graph. The set of instructions may be represented as a domain-specific language. The executable code may be executable code for a specific processor architecture.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system-on-a-chip (SoC), comprising:
 a central processing unit (CPU) to perform a compiler to generate code to accelerate matrix operations;   memory;   a Peripheral Component Interconnect (PCI) communication bus; and   a graphics processing unit (GPU) including a general processing cluster (GPC), where the GPC includes streaming multiprocessors (SMs) comprising:
 an instruction cache; 
 a dispatch unit; 
 cores; 
 a load/store unit (LSU); 
 shared memory; 
 an L1 cache; and 
   wherein the compiler is to:
 obtain a computer program; 
 extract a polyhedral representation of the computer program; 
 determine a transformation schedule using the polyhedral representation; and 
 generate executable code based on the transformation schedule and a processor architecture. 
   
     
     
         2 . The SoC of  claim 1 , wherein the GPU further comprises a scheduler unit. 
     
     
         3 . The SoC of  claim 1 , wherein the compiler is to further obtain one or more configuration files comprising parameters to help determine the transformation schedule. 
     
     
         4 . The SoC of  claim 1 , wherein the polyhedral representation of the computer program is a directed acyclic graph (DAG). 
     
     
         5 . The SoC of  claim 1 , wherein the GPU further comprises a memory partition unit. 
     
     
         6 . The SoC of  claim 1 , wherein the GPU further comprises a crossbar (Xbar). 
     
     
         7 . The SoC of  claim 1 , wherein the GPU further comprises an input/output (I/O) unit to interface with the PCI communication bus. 
     
     
         8 . The SoC of  claim 1 , further comprising a hub to interface with one or more GPU interconnects. 
     
     
         9 . The SoC of  claim 1 , wherein the GPC further comprises a raster engine. 
     
     
         10 . The SoC of  claim 1 , wherein the SMs each further comprise one or more interconnects. 
     
     
         11 . A system, comprising:
 a central processing unit (CPU) to perform a compiler to generate code to accelerate matrix operations;   memory;   a Peripheral Component Interconnect (PCI) communication bus; and   a graphics processing unit (GPU) including a general processing cluster (GPC), where the GPC includes streaming multiprocessors (SMs) comprising:
 an instruction cache; 
 a dispatch unit; 
 cores; 
 a load/store unit (LSU); 
 shared memory; 
 an L1 cache; and 
   wherein the compiler is to:
 obtain a computer program; 
 extract a polyhedral representation of the computer program; 
 determine a transformation schedule using the polyhedral representation; and 
 generate executable code based on the transformation schedule and a processor architecture. 
   
     
     
         12 . The system of  claim 11 , wherein the GPU further comprises a scheduler unit. 
     
     
         13 . The system of  claim 11 , wherein the SMs further comprise one or more special function units (SFUs). 
     
     
         14 . The system of  claim 11 , wherein the SMs further comprise a register file. 
     
     
         15 . The system of  claim 11 , further comprising one or more display devices. 
     
     
         16 . The system of  claim 11 , further comprising a network interface. 
     
     
         17 . The system of  claim 11 , further comprising a hub to interface with one or more GPU interconnects. 
     
     
         18 . The system of  claim 11 , wherein the GPC further comprises a raster engine. 
     
     
         19 . The system of  claim 11 , wherein the SMs each further comprise one or more interconnects. 
     
     
         20 . The system of  claim 11 , wherein the compiler is to further obtain one or more configuration files comprising parameters to help determine the transformation schedule.

Join the waitlist — get patent alerts

Track US2023251861A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.