Parallel slice processing method using a recirculating load-store queue for fast deallocation of issue queue entries
Abstract
A method of operation of a processor core execution unit circuit provides efficient use of area and energy by reducing the per-entry storage requirement of a load-store unit issue queue. The execution unit circuit includes a recirculation queue that stores the effective address of the load and store operations and the values to be stored by the store operations. A queue control logic controls the recirculation queue and issue queue so that that after the effective address of a load or store operation has been computed, the effective address of the load operation or the store operation is written to the recirculation queue and the operation is removed from the issue queue, so that address operands and other values that were in the issue queue entry no longer require storage. When a load or store operation is rejected by the cache unit, it is subsequently reissued from the recirculation queue.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of executing program instructions within a processor core, the method comprising:
receiving a stream of instructions including functional operations and load-store operations at an issue queue; computing effective addresses of load operations and store operations; issuing the load operations and store operations to a cache unit; storing entries corresponding to the load operations and the store operations at a recirculation queue; removing the load operations and store operations from the issue queue; and subsequently reissuing one of the load operations or store operations to the cache unit from the recirculation queue if the one of the load operations or store operations is rejected by the cache unit.
2 . The method of claim 1 , wherein the storing entries stores only the effective address of the load operations or store operations and for store operations, the value to be stored by the store operation.
3 . The method of claim 2 , further comprising:
removing load operations from the issue queue once the effective address is written to the recirculation queue; and removing store operations from the issue queue once the effective address and the values to be stored by the store operations are written to the recirculation queue.
4 . The method of claim 1 , further comprising:
removing load operations from the issue queue once the effective address is written to the recirculation queue; and issuing the store operations and the values to be stored by the store operations to the cache unit before removing the store data from the issue queue.
5 . The method of claim 1 , wherein the issuing issues the load and store operations to the cache unit in the same processor cycle as the storing entries stores the effective address of the load or store operation in the recirculation queue.
6 . The method of claim 1 , wherein the cache unit is implemented as a plurality of cache slices to which the load and store operations may be routed via a bus, and wherein the reissuing of the load operations or the store operations is directed to a different cache slice than another cache slice that has previously rejected the load operations or the store operations.Join the waitlist — get patent alerts
Track US2016202988A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.