US2019304049A1PendingUtilityA1

Dynamic thread execution arbitration

Assignee: INTEL CORPPriority: Apr 21, 2017Filed: Mar 1, 2019Published: Oct 3, 2019
Est. expiryApr 21, 2037(~10.7 yrs left)· nominal 20-yr term from priority
G06F 9/4881G06T 2210/52G06T 1/20G06F 9/4831G06F 9/505
58
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A mechanism is described for facilitating thread execution arbitration for thread scheduling relating to graphics processors at computing devices. A method of embodiments, as described herein, includes assigning priority levels to threads based on stall signals communicated from the one or more shared function units to one or more execution units of a processor including a graphics processor, and selecting a first thread to be scheduled and a second thread to be ignored based on the stall signals.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled) 
     
     
         21 . An apparatus comprising:
 one or more processors including a graphics processor, the one or more processors including a plurality of processing units, a plurality of thread schedulers to schedule threads for the plurality of processing units, and a processing pipeline; and   one or more shared function units, each shared function unit being shared by two or more of the plurality of processing units of the one or more processors;   wherein the one or more processors are to:
 detect and observe stall signals communicated from the one or more shared function units to the processing units of the one or more processors, 
 predict one or more stalls to occur in the processing pipeline based on the received stall signals, 
 estimate an effect on threads in a plurality of threads to be caused by the predicted one or more stalls in the processing pipeline, and 
 schedule threads of the plurality of threads based at least in part on the predicted effect on threads in the plurality of threads. 
   
     
     
         22 . The apparatus of  claim 21 , wherein scheduling threads of the plurality of threads includes the one or more processors to select a first thread of the plurality of threads to be scheduled and a second thread of the plurality of threads to be ignored based at least in part on the predicted effect on threads of the plurality of threads. 
     
     
         23 . The apparatus of  claim 21 , wherein the one or more processors are further to assign priority levels to the plurality of threads based at least in part on the predicted one or more stalls in the processing pipeline. 
     
     
         24 . The apparatus of  claim 23 , wherein the assignment of priority levels includes determining whether a thread is deserving of a high priority assignment or a low priority assignment based at least in part on the predicted one or more stalls in the processing pipeline. 
     
     
         25 . The apparatus of  claim 24 , wherein the high priority and low priority assignments are further determined based on one or more predetermined thresholds indicating a level of stall that is acceptable or unacceptable for scheduling of the threads. 
     
     
         26 . The apparatus of  claim 21 , wherein the plurality of processing units includes a plurality of streaming multiprocessors (SMs). 
     
     
         27 . The apparatus of  claim 21 , wherein the shared function units include one or more of a sampler, a data port, a shared local memory, and a pixel/color pipe. 
     
     
         28 . The apparatus of  claim 21 , wherein the graphics processor is co-located with an application processor on a common semiconductor package. 
     
     
         29 . A method comprising:
 detecting and observing stall signals communicated from one or more shared function units to a plurality of processing units of one or more processors including a graphics processor, each shared function unit being shared by two or more of the plurality of processing units of the one or more processors;   predicting one or more stalls to occur in a processing pipeline based on the received stall signals;   estimating an effect on threads in a plurality of threads to be caused by the predicted one or more stalls in the processing pipeline; and   scheduling threads of the plurality of threads based at least in part on the predicted effect on threads in the plurality of threads.   
     
     
         30 . The method of  claim 29 , wherein scheduling threads of the plurality of threads includes selecting a first thread of the plurality of threads to be scheduled and a second thread of the plurality of threads to be ignored based at least in part on the predicted effect on threads of the plurality of threads. 
     
     
         31 . The method of  claim 29 , further comprising
 assigning priority levels to the plurality of threads based at least in part on the predicted one or more stalls in the processing pipeline.   
     
     
         32 . The method of  claim 31 , wherein the assignment of priority levels includes determining whether a thread is deserving of a high priority assignment or a low priority assignment based at least in part on the predicted one or more stalls in the processing pipeline. 
     
     
         33 . The method of  claim 32 , wherein the high priority and low priority assignments are further determined based on one or more predetermined thresholds indicating a level of stall that is acceptable or unacceptable for scheduling of the threads. 
     
     
         34 . The method of  claim 29 , wherein the plurality of processing units includes a plurality of streaming multiprocessors (SMs). 
     
     
         35 . The method of  claim 29 , wherein the shared function units include one or more of a sampler, a data port, a shared local memory, and a pixel/color pipe. 
     
     
         36 . At least one non-transitory machine-readable medium comprising instructions that when executed by a computing device, cause the computing device to perform operations comprising:
 detecting and observing stall signals communicated from one or more shared function units to a plurality of processing units of one or more processors including a graphics processor, each shared function unit being shared by two or more of the plurality of processing units of the one or more processors;   predicting one or more stalls to occur in a processing pipeline based on the received stall signals;   estimating an effect on threads in a plurality of threads to be caused by the predicted one or more stalls in the processing pipeline; and   scheduling threads of the plurality of threads based at least in part on the predicted effect on threads in the plurality of threads.   
     
     
         37 . The machine-readable medium of  claim 36 , wherein scheduling threads of the plurality of threads includes selecting a first thread of the plurality of threads to be scheduled and a second thread of the plurality of threads to be ignored based at least in part on the predicted effect on threads of the plurality of threads. 
     
     
         38 . The machine-readable medium of  claim 36 , further comprising instructions that, when executed by the computing device, cause the computing device to perform operations comprising:
 assigning priority levels to the plurality of threads based at least in part on the predicted one or more stalls in the processing pipeline.   
     
     
         39 . The machine-readable medium of  claim 38 , wherein the assignment of priority levels includes determining whether a thread is deserving of a high priority assignment or a low priority assignment based at least in part on the predicted one or more stalls in the processing pipeline. 
     
     
         40 . The machine-readable medium of  claim 39 , wherein the high priority and low priority assignments are further determined based on one or more predetermined thresholds indicating a level of stall that is acceptable or unacceptable for scheduling of the threads. 
     
     
         41 . The machine-readable medium of  claim 36 , wherein the plurality of processing units includes a plurality of streaming multiprocessors (SMs). 
     
     
         42 . The machine-readable medium of  claim 36 , wherein the shared function units include one or more of a sampler, a data port, a shared local memory, and a pixel/color pipe.

Join the waitlist — get patent alerts

Track US2019304049A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.