US2006015547A1PendingUtilityA1

Efficient circuits for out-of-order microprocessors

Assignee: UNIV YALEPriority: Mar 12, 1998Filed: Mar 7, 2005Published: Jan 19, 2006
Est. expiryMar 12, 2018(expired)· nominal 20-yr term from priority
G06F 9/3867G06F 9/3802G06F 9/3885G06F 2207/5063G06F 7/506G06F 9/3869G06F 9/3836G06F 9/384G06F 9/3844G06F 9/3838G06F 9/3854G06F 9/3856G06F 9/3858
44
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The poor scalability of existing superscalar processors has been of great concern to the computer engineering community. In particular, the critical-path delays of many components in existing implementations grow quadratically with the issue width and the window size. This patent presents a novel way to reimplement these components and reduce their critical-path delay growth. It then describes an entire processor microarchitecture, called the Ultrascalar processor, that has better critical-path delay growth than existing superscalars. Most of our scalable designs are based on a single circuit, a cyclic segmented parallel prefix (cspp). We observe that processor components typically operate on a wrap-around sequence of instructions, computing some associative property of that sequence. For example, to assign an ALU to the oldest requesting instruction, each instruction in the instruction sequence must be told whether any preceding instructions are requesting an ALU. Similarly, to read an argument register, an instruction must somehow communicate with the most recent preceding instruction that wrote that register. A cspp circuit can implement such functions by computing for each instruction within a wrap-around instruction sequence the accumulative result of applying some associative operator to all the preceding instructions. A cspp circuit has a critical path gate delay logarithmic in the length of the instruction sequence. Depending on its associative operation and its layout, a cspp circuit can have a critical path wire delay sublinear in the length of the instruction sequence.

Claims

exact text as granted — not AI-modified
1 - 2 . (canceled)  
   
   
       3 . An ultrascalar processor comprising: 
 a fetch stage;    a rename stage for processing output from the fetch stage;    an analyze stage for processing output from the rename stage;    a schedule stage for processing output from the analyze stage;    an execute stage for processing output from the schedule; and    a broadcast stage for processing output from the execute stage to the analyze stage,    wherein the scheduler is implemented using a cyclic segmented parallel prefix circuit.    
   
   
       4 . An ultrascalar processor comprising: 
 a fetch stage;    a rename stage for processing output from the fetch stage;    an analyze stage for processing output from the rename stage;    a schedule stage for processing output from the analyze stage;    an execute stage for processing output from the schedule; and    a broadcast stage for processing output from the execute stage to the analyze stage,    wherein a parallel-prefix summation circuit computes program counts in the fetch stage.

Join the waitlist — get patent alerts

Track US2006015547A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.