US2019304049A1PendingUtilityA1
Dynamic thread execution arbitration
Est. expiryApr 21, 2037(~10.7 yrs left)· nominal 20-yr term from priority
Inventors:Joydeep RayAbhishek R. AppuSubramaniam MaiyuranEric J. HoekstraPrasoonkumar SurtiBalaji VembuAltug Koker
G06F 9/4881G06T 2210/52G06T 1/20G06F 9/4831G06F 9/505
58
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A mechanism is described for facilitating thread execution arbitration for thread scheduling relating to graphics processors at computing devices. A method of embodiments, as described herein, includes assigning priority levels to threads based on stall signals communicated from the one or more shared function units to one or more execution units of a processor including a graphics processor, and selecting a first thread to be scheduled and a second thread to be ignored based on the stall signals.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . An apparatus comprising:
one or more processors including a graphics processor, the one or more processors including a plurality of processing units, a plurality of thread schedulers to schedule threads for the plurality of processing units, and a processing pipeline; and one or more shared function units, each shared function unit being shared by two or more of the plurality of processing units of the one or more processors; wherein the one or more processors are to:
detect and observe stall signals communicated from the one or more shared function units to the processing units of the one or more processors,
predict one or more stalls to occur in the processing pipeline based on the received stall signals,
estimate an effect on threads in a plurality of threads to be caused by the predicted one or more stalls in the processing pipeline, and
schedule threads of the plurality of threads based at least in part on the predicted effect on threads in the plurality of threads.
22 . The apparatus of claim 21 , wherein scheduling threads of the plurality of threads includes the one or more processors to select a first thread of the plurality of threads to be scheduled and a second thread of the plurality of threads to be ignored based at least in part on the predicted effect on threads of the plurality of threads.
23 . The apparatus of claim 21 , wherein the one or more processors are further to assign priority levels to the plurality of threads based at least in part on the predicted one or more stalls in the processing pipeline.
24 . The apparatus of claim 23 , wherein the assignment of priority levels includes determining whether a thread is deserving of a high priority assignment or a low priority assignment based at least in part on the predicted one or more stalls in the processing pipeline.
25 . The apparatus of claim 24 , wherein the high priority and low priority assignments are further determined based on one or more predetermined thresholds indicating a level of stall that is acceptable or unacceptable for scheduling of the threads.
26 . The apparatus of claim 21 , wherein the plurality of processing units includes a plurality of streaming multiprocessors (SMs).
27 . The apparatus of claim 21 , wherein the shared function units include one or more of a sampler, a data port, a shared local memory, and a pixel/color pipe.
28 . The apparatus of claim 21 , wherein the graphics processor is co-located with an application processor on a common semiconductor package.
29 . A method comprising:
detecting and observing stall signals communicated from one or more shared function units to a plurality of processing units of one or more processors including a graphics processor, each shared function unit being shared by two or more of the plurality of processing units of the one or more processors; predicting one or more stalls to occur in a processing pipeline based on the received stall signals; estimating an effect on threads in a plurality of threads to be caused by the predicted one or more stalls in the processing pipeline; and scheduling threads of the plurality of threads based at least in part on the predicted effect on threads in the plurality of threads.
30 . The method of claim 29 , wherein scheduling threads of the plurality of threads includes selecting a first thread of the plurality of threads to be scheduled and a second thread of the plurality of threads to be ignored based at least in part on the predicted effect on threads of the plurality of threads.
31 . The method of claim 29 , further comprising
assigning priority levels to the plurality of threads based at least in part on the predicted one or more stalls in the processing pipeline.
32 . The method of claim 31 , wherein the assignment of priority levels includes determining whether a thread is deserving of a high priority assignment or a low priority assignment based at least in part on the predicted one or more stalls in the processing pipeline.
33 . The method of claim 32 , wherein the high priority and low priority assignments are further determined based on one or more predetermined thresholds indicating a level of stall that is acceptable or unacceptable for scheduling of the threads.
34 . The method of claim 29 , wherein the plurality of processing units includes a plurality of streaming multiprocessors (SMs).
35 . The method of claim 29 , wherein the shared function units include one or more of a sampler, a data port, a shared local memory, and a pixel/color pipe.
36 . At least one non-transitory machine-readable medium comprising instructions that when executed by a computing device, cause the computing device to perform operations comprising:
detecting and observing stall signals communicated from one or more shared function units to a plurality of processing units of one or more processors including a graphics processor, each shared function unit being shared by two or more of the plurality of processing units of the one or more processors; predicting one or more stalls to occur in a processing pipeline based on the received stall signals; estimating an effect on threads in a plurality of threads to be caused by the predicted one or more stalls in the processing pipeline; and scheduling threads of the plurality of threads based at least in part on the predicted effect on threads in the plurality of threads.
37 . The machine-readable medium of claim 36 , wherein scheduling threads of the plurality of threads includes selecting a first thread of the plurality of threads to be scheduled and a second thread of the plurality of threads to be ignored based at least in part on the predicted effect on threads of the plurality of threads.
38 . The machine-readable medium of claim 36 , further comprising instructions that, when executed by the computing device, cause the computing device to perform operations comprising:
assigning priority levels to the plurality of threads based at least in part on the predicted one or more stalls in the processing pipeline.
39 . The machine-readable medium of claim 38 , wherein the assignment of priority levels includes determining whether a thread is deserving of a high priority assignment or a low priority assignment based at least in part on the predicted one or more stalls in the processing pipeline.
40 . The machine-readable medium of claim 39 , wherein the high priority and low priority assignments are further determined based on one or more predetermined thresholds indicating a level of stall that is acceptable or unacceptable for scheduling of the threads.
41 . The machine-readable medium of claim 36 , wherein the plurality of processing units includes a plurality of streaming multiprocessors (SMs).
42 . The machine-readable medium of claim 36 , wherein the shared function units include one or more of a sampler, a data port, a shared local memory, and a pixel/color pipe.Join the waitlist — get patent alerts
Track US2019304049A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.