US2010122067A1PendingUtilityA1

Across-thread out-of-order instruction dispatch in a multithreaded microprocessor

Assignee: NVIDIA CORPPriority: Dec 18, 2003Filed: Jan 20, 2010Published: May 13, 2010
Est. expiryDec 18, 2023(expired)· nominal 20-yr term from priority
G06F 9/3888G06F 9/3851G06F 9/3802
49
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Instruction dispatch in a multithreaded microprocessor such as a graphics processor is not constrained by an order among the threads. Instructions for each thread are fetched, and a dispatch circuit determines which instructions in the buffer are ready to execute. The dispatch circuit may issue any ready instruction for execution, and an instruction from one thread may be issued prior to an instruction from another thread regardless of which instruction was fetched first. If multiple functional units are available, multiple instructions can be dispatched in parallel.

Claims

exact text as granted — not AI-modified
1 .- 20 . (canceled) 
   
   
       21 . A processor configured for parallel processing of a plurality of threads, the processor comprising:
 a priority encoder configured to assigning a priority ranking to each of a plurality of different thread types of the plurality of threads;   a fetch module configured to:
 fetch a first instruction for a first one of the plurality of threads; and 
 fetch, after the first instruction is fetched, a plurality of additional instructions for additional ones of the plurality of threads; and 
   a dispatch module, communicatively coupled with the priority encoder and the fetch module, and configured to:
 issue, during a latency period associated with the first thread, a second instruction of the plurality of additional instructions, the second instruction selected for issue based at least in part on the priority ranking associated with the second instruction's thread and an amount of time the second instruction has been ready to issue; and 
 issue the first instruction after the second instruction is issued. 
   
   
   
       22 . The processor of  claim 21 , wherein the second instruction's thread is of a first type assigned a priority ranking higher than priority rankings of each other of the additional plurality of threads. 
   
   
       23 . The processor of  claim 21 , wherein the second instruction's thread is of a first type different from types associated with each other of the additional plurality of threads. 
   
   
       24 . The processor of  claim 23 , wherein the fetch module fetches the second instruction after fetching each other of the additional plurality of threads. 
   
   
       25 . The processor of  claim 21 , wherein the plurality of different thread types are defined by a type of input data each thread type processes. 
   
   
       26 . The processor of  claim 25 , wherein a first thread type of the plurality of different thread types and a second thread type of the plurality of different thread types each executes a same program on different input data. 
   
   
       27 . The processor of  claim 25 , wherein the plurality of different thread types comprise:
 a vertex thread type assigned a first priority ranking; and   a pixel thread type assigned a second priority ranking.   
   
   
       28 . The processor of  claim 21 , wherein the plurality of different thread types are defined by a program each thread type processes. 
   
   
       29 . The processor of  claim 21 , wherein the first instruction and the second instruction are instructions from different portions of a same program. 
   
   
       30 . The processor of  claim 21 , wherein the fetch module is further configured to fetch instructions from the plurality of threads according to a priority ranking for fetching thread types. 
   
   
       31 . The processor of  claim 30 , wherein the priority ranking for fetching thread types comprises the priority ranking used by the dispatch module. 
   
   
       32 . The processor of  claim 21 , wherein the processor comprises a graphics processor. 
   
   
       33 . A method for executing a plurality of threads in a multithreaded processor, the method comprising:
 assigning a priority ranking to each of a plurality of different thread types of the plurality of threads;   fetching a first instruction for a first one of the plurality of threads;   fetching, after the first instruction is fetched, a plurality of additional instructions for additional ones of the plurality of threads;   during a latency period associated with the first thread, issuing a second instruction of the plurality of additional instructions, the second instruction selected for issue based at least in part on the priority ranking associated with the second instruction's thread and an amount of time the second instruction has been ready to issue; and   issuing the first instruction after the second instruction is issued.   
   
   
       34 . The method of  claim 33 , wherein the second instruction's thread is of a first type assigned a priority ranking higher than priority rankings for types associated with each other of the additional plurality of threads. 
   
   
       35 . The method of  claim 34 , wherein a fetch module fetches the second instruction after fetching each other of the additional plurality of threads. 
   
   
       36 . The method of  claim 33 , wherein the plurality of different thread types are defined by a different type of input data each thread type processes. 
   
   
       37 . The method of  claim 36 , wherein the plurality of different thread types comprise:
 a vertex thread type assigned a first priority ranking; and   a pixel thread type assigned a second priority ranking indicating less importance than the first priority ranking.   
   
   
       38 . The method of  claim 33 , wherein the plurality of different thread types are defined by a different program each thread type processes. 
   
   
       39 . The method of  claim 33 , wherein the multithreaded processor comprises a graphics processor. 
   
   
       40 . A device for executing a plurality of threads in a multithreaded processor, the device comprising:
 means for assigning a priority ranking to each of a plurality of different thread types of the plurality of threads;   means for fetching a first instruction for a first one of the plurality of threads;   means for fetching, after the first instruction is fetched, a plurality of additional instructions for additional ones of the plurality of threads;   means for issuing, during a latency period associated with the first thread wherein the first instruction fails to become ready to issue, a second instruction of the plurality of additional instructions, the second instruction selected for issue based at least in part on the priority ranking associated with the second instruction's thread and an amount of time the second instruction has been ready to issue; and   means for issuing the first instruction after the second instruction is issued.

Join the waitlist — get patent alerts

Track US2010122067A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.