Thread group scheduling for graphics processing
Abstract
Embodiments are generally directed to thread group scheduling for graphics processing. An embodiment of an apparatus includes a plurality of processors including a plurality of graphics processors to process data; a memory; and one or more caches for storage of data for the plurality of graphics processors, wherein the one or more processors are to schedule a plurality of groups of threads for processing by the plurality of graphics processors, the scheduling of the plurality of groups of threads including the plurality of processors to apply a bias for scheduling the plurality of groups of threads according to a cache locality for the one or more caches.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . An apparatus comprising:
one or more processors including one or more graphics processors to process data, the one or more processors including a scheduler to schedule a plurality of thread groups for processing; and a memory to store data, the data including data for the plurality of thread groups scheduled for processing; wherein the scheduler is to schedule a plurality of warps for processing, each of the plurality of warps including a plurality of thread groups, scheduling the plurality of warps including assigning the thread groups of the plurality of warps to a plurality of sub-blocks, each sub-block being included in a one of a plurality of blocks representing assignment of thread groups to one or more graphics processors of the one or more graphics processors; wherein assigning the thread groups of the plurality of warps includes assigning the thread groups of each of the plurality of warps to a respective sub-block of the plurality of sub-blocks.
22 . The apparatus of claim 21 , wherein scheduling the thread groups of the plurality of warps includes scheduling the thread groups according to the assignment of the thread groups to the plurality of sub-blocks.
23 . The apparatus of claim 22 , wherein scheduling the thread groups of the plurality of warps according to the assignment of the thread groups to the plurality of sub-blocks includes scheduling the thread groups to operate synchronously across multiple warps of the plurality of warps.
24 . The apparatus of claim 22 , wherein scheduling the thread groups of the plurality of warps according to the assignment of the thread groups to the plurality of sub-blocks includes scheduling the thread groups of the plurality of warps to operate sequentially.
25 . The apparatus of claim 22 , wherein each warp is bound to the respective sub-block to which the thread groups of the warp are assigned.
26 . The apparatus of claim 22 , wherein, upon the thread groups of the plurality of warps diverging from each other, scheduler is to schedule the thread groups of each warp independently of the thread groups of other warps.
27 . The apparatus of claim 21 , further comprising one or more caches for storage of data for the plurality of graphics processors, the one or more caches including storage of data for the thread groups scheduled for processing.
28 . One or more non-transitory computer-readable storage mediums having stored thereon executable computer program instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:
identifying a plurality of warps for scheduling for processing by one or more graphics processors, each of the plurality of warps including a plurality of thread groups; and assigning thread groups of the plurality of warps to a plurality of sub-blocks, each sub-block being included in a one of a plurality of blocks representing assignment of thread groups to one or more graphics processors; wherein assigning the thread groups of the plurality of warps includes assigning the thread groups of each of the plurality of warps to a respective sub-block of the plurality of sub-blocks.
29 . The one or more computer-readable storage mediums of claim 28 , wherein scheduling the thread groups of the plurality of warps includes scheduling the thread groups according to the assignment of the thread groups to the plurality of sub-blocks.
30 . The one or more computer-readable storage mediums of claim 29 , wherein scheduling the thread groups of the plurality of warps according to the assignment of the thread groups to the plurality of sub-blocks includes scheduling the thread groups to operate synchronously across multiple warps of the plurality of warps.
31 . The one or more computer-readable storage mediums of claim 29 , wherein scheduling the thread groups of the plurality of warps according to the assignment of the thread groups to the plurality of sub-blocks includes scheduling the thread groups of the plurality of warps to operate sequentially.
32 . The one or more computer-readable storage mediums of claim 28 , wherein each warp is bound to the respective sub-block to which the thread groups of the warp are assigned.
33 . The one or more computer-readable storage mediums of claim 28 , the instructions further including instructions for:
upon the thread groups of the plurality of warps diverging from each other, scheduling the thread groups of each warp independently of the thread groups of other warps.
34 . A method comprising:
identifying a plurality of warps for scheduling for processing by one or more graphics processors, each of the plurality of warps including a plurality of thread groups; and assigning thread groups of the plurality of warps to a plurality of sub-blocks, each sub-block being included in a one of a plurality of blocks representing assignment of thread groups to one or more graphics processors; wherein assigning the thread groups of the plurality of warps includes assigning the thread groups of each of the plurality of warps to a respective sub-block of the plurality of sub-blocks.
35 . The method of claim 34 , wherein scheduling the thread groups of the plurality of warps includes scheduling the thread groups according to the assignment of the thread groups to the plurality of sub-blocks.
36 . The method of claim 35 , wherein scheduling the thread groups of the plurality of warps according to the assignment of the thread groups to the plurality of sub-blocks includes scheduling the thread groups to operate synchronously across multiple warps of the plurality of warps.
37 . The method of claim 35 , wherein scheduling the thread groups of the plurality of warps according to the assignment of the thread groups to the plurality of sub-blocks includes scheduling the thread groups of the plurality of warps to operate sequentially.
38 . The method of claim 34 , wherein each warp is bound to the respective sub-block to which the thread groups of the warp are assigned.
39 . The method of claim 34 , further comprising:
upon the thread groups of the plurality of warps diverging from each other, scheduling the thread groups of each warp independently of the thread groups of other warps.Join the waitlist — get patent alerts
Track US2024028404A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.