US2020371804A1PendingUtilityA1
Boosting local memory performance in processor graphics
Est. expiryOct 29, 2035(~9.3 yrs left)· nominal 20-yr term from priority
Inventors:Can K. Que
G06T 1/20G06F 9/3887G06F 9/3888G06F 9/38885G06F 12/0804G06F 2212/1024G06F 9/3867G06T 19/00G06F 9/30138G06F 9/522G06F 2212/1016G06T 1/60
51
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
In some cases, processor graphics with a slower local memory can compensate by using another memory in place of the lowest level or L3 cache. For example, in some processors, there is a large register space that can be used for the local memory function by allocating the local memory within those registers. Also, since the registers do not operate with barriers, barriers can be simulated by letting one execution unit thread execute more SIMD instructions. For example, one execution thread may simulate a whole work-group in the OpenCL API.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
executing a first work-group on a processor, wherein the first work-group comprises a plurality of work-items that are executed in parallel to perform a defined function, wherein the processor comprises internal registers files, wherein the processor uses a cache hierarchy including a lowest level cache and at least one other cache; determining an available space in the internal register files of the processor; and allocating the determined available space in the internal register files as local memory for the first work-group, wherein the local memory for the first work-group is only accessible by the plurality of work-items included in the first work-group.
2 . The method of claim 1 including:
determining whether the available space in the internal register files is sufficient for the local memory for the first work-group; and
in response to a determination that the available space in the internal register files is not sufficient for the local memory for the first work-group:
allocating the available space in the internal register files as a first portion of the local memory for the first work-group, and
allocating the lowest level cache as a second portion of the local memory for the first work-group.
3 . The method of claim 1 wherein the first work-group is an OpenCL work-group.
4 . The method of claim 1 including:
detecting a non-aligned write from a first register location in a source register to a second register location in a destination register, wherein the first register location is not aligned with the second register location;
in response to the detected non-aligned write, generating a plurality of instructions based on permutation group theory; and
transforming the non-aligned write using the generated plurality of instructions.
5 . The method of claim 4 including generating the plurality of instructions in an OpenGL compiler.
6 . The method of claim 4 wherein generating the plurality of instructions comprises generating log (k−1) instructions to transform the source register, wherein k is a value calculated using permutation group theory.
7 . The method of claim 6 , wherein k is a value calculated using permutation group theory.
8 . A non-transitory computer readable medium storing instructions, the instructions executable by a processor to:
execute a first work-group on a processor, wherein the first work-group comprises a plurality of work-items that are executed in parallel to perform a defined function, wherein the processor comprises internal registers files, wherein the processor uses a cache hierarchy including a lowest level cache and at least one other cache; determine an available space in the internal register files of the processor; and allocate the determined available space in the internal register files as local memory for the first work-group, wherein the local memory for the first work-group is only accessible by the plurality of work-items included in the first work-group.
9 . The non-transitory computer readable medium of claim 8 , wherein the instructions are executable by the processor to:
determine whether the available space in the internal register files is sufficient for the local memory for the first work-group; and in response to a determination that the available space in the internal register files is not sufficient for the local memory for the first work-group:
allocate the available space in the internal register files as a first portion of the local memory for the first work-group, and
allocate the lowest level cache as a second portion of the local memory for the first work-group.
10 . The non-transitory computer readable medium of claim 8 , wherein the first work-group is an OpenCL work-group.
11 . The non-transitory computer readable medium of claim 8 , wherein the instructions are executable by the processor to:
detect a non-aligned write from a first register location in a source register to a second register location in a destination register, wherein the first register location is not aligned with the second register location; in response to the detected non-aligned write, generate a plurality of instructions based on permutation group theory; and transform the non-aligned write using the generated plurality of instructions.
12 . The non-transitory computer readable medium of claim 11 , wherein the instructions are executable by the processor to:
generate the plurality of instructions in an OpenGL compiler
13 . The non-transitory computer readable medium of claim 11 , wherein generating the plurality of instructions comprises generating log (k−1) instructions to transform the source register.
14 . The non-transitory computer readable medium of claim 13 , wherein k is a value calculated using permutation group theory.
15 . A system comprising:
a processor to:
execute a first work-group comprising a plurality of work-items that are executed in parallel to perform a defined function, wherein the processor comprises internal registers files, wherein the processor uses a cache hierarchy including a lowest level cache and at least one other cache;
determine an available space in the internal register files of the processor; and
allocate the determined available space in the internal register files as local memory for the first work-group, wherein the local memory for the first work-group is only accessible by the plurality of work-items included in the first work-group; and
a memory operatively coupled to the processor.
16 . The system of claim 15 , the processor to:
determine whether the available space in the internal register files is sufficient for the local memory for the first work-group; and in response to a determination that the available space in the internal register files is not sufficient for the local memory for the first work-group:
allocate the available space in the internal register files as a first portion of the local memory for the first work-group, and
allocate the lowest level cache as a second portion of the local memory for the first work-group.
17 . The system of claim 15 , wherein the first work-group is an OpenCL work-group.
18 . The system of claim 15 , the processor to:
detect a non-aligned write from a first register location in a source register to a second register location in a destination register, wherein the first register location is not aligned with the second register location; in response to the detected non-aligned write, generate a plurality of instructions based on permutation group theory; and transform the non-aligned write using the generated plurality of instructions.
19 . The system of claim 15 , the processor to:
generate the plurality of instructions in an OpenGL compiler
20 . The system of claim 15 , wherein generating the plurality of instructions comprises generating log (k−1) instructions to transform the source register, wherein k is a value calculated using permutation group theory.Join the waitlist — get patent alerts
Track US2020371804A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.