US2020371804A1PendingUtilityA1

Boosting local memory performance in processor graphics

Assignee: INTEL CORPPriority: Oct 29, 2015Filed: Aug 6, 2020Published: Nov 26, 2020
Est. expiryOct 29, 2035(~9.3 yrs left)· nominal 20-yr term from priority
Inventors:Can K. Que
G06T 1/20G06F 9/3887G06F 9/3888G06F 9/38885G06F 12/0804G06F 2212/1024G06F 9/3867G06T 19/00G06F 9/30138G06F 9/522G06F 2212/1016G06T 1/60
51
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In some cases, processor graphics with a slower local memory can compensate by using another memory in place of the lowest level or L3 cache. For example, in some processors, there is a large register space that can be used for the local memory function by allocating the local memory within those registers. Also, since the registers do not operate with barriers, barriers can be simulated by letting one execution unit thread execute more SIMD instructions. For example, one execution thread may simulate a whole work-group in the OpenCL API.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 executing a first work-group on a processor, wherein the first work-group comprises a plurality of work-items that are executed in parallel to perform a defined function, wherein the processor comprises internal registers files, wherein the processor uses a cache hierarchy including a lowest level cache and at least one other cache;   determining an available space in the internal register files of the processor; and   allocating the determined available space in the internal register files as local memory for the first work-group, wherein the local memory for the first work-group is only accessible by the plurality of work-items included in the first work-group.   
     
     
         2 . The method of  claim 1  including:
 determining whether the available space in the internal register files is sufficient for the local memory for the first work-group; and 
 in response to a determination that the available space in the internal register files is not sufficient for the local memory for the first work-group:
 allocating the available space in the internal register files as a first portion of the local memory for the first work-group, and 
 allocating the lowest level cache as a second portion of the local memory for the first work-group. 
 
 
     
     
         3 . The method of  claim 1  wherein the first work-group is an OpenCL work-group. 
     
     
         4 . The method of  claim 1  including:
 detecting a non-aligned write from a first register location in a source register to a second register location in a destination register, wherein the first register location is not aligned with the second register location; 
 in response to the detected non-aligned write, generating a plurality of instructions based on permutation group theory; and 
 transforming the non-aligned write using the generated plurality of instructions. 
 
     
     
         5 . The method of  claim 4  including generating the plurality of instructions in an OpenGL compiler. 
     
     
         6 . The method of  claim 4  wherein generating the plurality of instructions comprises generating log (k−1) instructions to transform the source register, wherein k is a value calculated using permutation group theory. 
     
     
         7 . The method of  claim 6 , wherein k is a value calculated using permutation group theory. 
     
     
         8 . A non-transitory computer readable medium storing instructions, the instructions executable by a processor to:
 execute a first work-group on a processor, wherein the first work-group comprises a plurality of work-items that are executed in parallel to perform a defined function, wherein the processor comprises internal registers files, wherein the processor uses a cache hierarchy including a lowest level cache and at least one other cache;   determine an available space in the internal register files of the processor; and   allocate the determined available space in the internal register files as local memory for the first work-group, wherein the local memory for the first work-group is only accessible by the plurality of work-items included in the first work-group.   
     
     
         9 . The non-transitory computer readable medium of  claim 8 , wherein the instructions are executable by the processor to:
 determine whether the available space in the internal register files is sufficient for the local memory for the first work-group; and   in response to a determination that the available space in the internal register files is not sufficient for the local memory for the first work-group:
 allocate the available space in the internal register files as a first portion of the local memory for the first work-group, and 
 allocate the lowest level cache as a second portion of the local memory for the first work-group. 
   
     
     
         10 . The non-transitory computer readable medium of  claim 8 , wherein the first work-group is an OpenCL work-group. 
     
     
         11 . The non-transitory computer readable medium of  claim 8 , wherein the instructions are executable by the processor to:
 detect a non-aligned write from a first register location in a source register to a second register location in a destination register, wherein the first register location is not aligned with the second register location;   in response to the detected non-aligned write, generate a plurality of instructions based on permutation group theory; and   transform the non-aligned write using the generated plurality of instructions.   
     
     
         12 . The non-transitory computer readable medium of  claim 11 , wherein the instructions are executable by the processor to:
 generate the plurality of instructions in an OpenGL compiler   
     
     
         13 . The non-transitory computer readable medium of  claim 11 , wherein generating the plurality of instructions comprises generating log (k−1) instructions to transform the source register. 
     
     
         14 . The non-transitory computer readable medium of  claim 13 , wherein k is a value calculated using permutation group theory. 
     
     
         15 . A system comprising:
 a processor to:
 execute a first work-group comprising a plurality of work-items that are executed in parallel to perform a defined function, wherein the processor comprises internal registers files, wherein the processor uses a cache hierarchy including a lowest level cache and at least one other cache; 
 determine an available space in the internal register files of the processor; and 
 allocate the determined available space in the internal register files as local memory for the first work-group, wherein the local memory for the first work-group is only accessible by the plurality of work-items included in the first work-group; and 
   a memory operatively coupled to the processor.   
     
     
         16 . The system of  claim 15 , the processor to:
 determine whether the available space in the internal register files is sufficient for the local memory for the first work-group; and   in response to a determination that the available space in the internal register files is not sufficient for the local memory for the first work-group:
 allocate the available space in the internal register files as a first portion of the local memory for the first work-group, and 
 allocate the lowest level cache as a second portion of the local memory for the first work-group. 
   
     
     
         17 . The system of  claim 15 , wherein the first work-group is an OpenCL work-group. 
     
     
         18 . The system of  claim 15 , the processor to:
 detect a non-aligned write from a first register location in a source register to a second register location in a destination register, wherein the first register location is not aligned with the second register location;   in response to the detected non-aligned write, generate a plurality of instructions based on permutation group theory; and   transform the non-aligned write using the generated plurality of instructions.   
     
     
         19 . The system of  claim 15 , the processor to:
 generate the plurality of instructions in an OpenGL compiler   
     
     
         20 . The system of  claim 15 , wherein generating the plurality of instructions comprises generating log (k−1) instructions to transform the source register, wherein k is a value calculated using permutation group theory.

Join the waitlist — get patent alerts

Track US2020371804A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.