Prefetching associated with predicated store instructions
Abstract
Technology related to prefetching data associated with predicated stores of programs in block-based processor architectures is disclosed. In one example of the disclosed technology, a processor includes a block-based processor core for executing an instruction block comprising a plurality of instructions. The block-based processor core includes decode logic and prefetch logic. The decode logic is configured to detect a predicated store instruction of the instruction block. The prefetch logic is configured to calculate a target address of the predicated store instruction and initiate a memory operation associated with the calculated target address before a predicate of the predicated store instruction is calculated.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A processor comprising a block-based processor core for executing an instruction block comprising an instruction header and a plurality of instructions, the block-based processor core comprising:
decode logic configured to detect a predicated store instruction of the instruction block; and prefetch logic configured to:
receive a first value associated with the predicated store instruction;
calculate a target address of the predicated store instruction using the received first value; and
initiate a memory operation associated with the calculated target address before a predicate of the predicated store instruction is calculated.
2 . The block-based processor core of claim 1 , wherein the memory operation includes issuing a prefetch request to a memory hierarchy of the processor to prefetch a cache line spanning the calculated target address.
3 . The block-based processor core of claim 1 , wherein the memory operation includes fetching coherence permissions for a memory line including data at the calculated target address.
4 . The block-based processor core of claim 1 , wherein the memory operation includes determining whether an inter-thread conflict exists for a memory line spanning the calculated target address.
5 . The block-based processor core of claim 1 , wherein the target address is calculated using a dedicated arithmetic unit of the prefetch logic.
6 . The block-based processor core of claim 1 , wherein the first value is generated by another instruction of the instruction block that targets the predicated store instruction.
7 . The block-based processor core of claim 1 , wherein calculating the target address comprises performing the target address calculation during an open instruction issue slot and using an arithmetic unit of instruction execution logic.
8 . The block-based processor core of claim 1 , wherein the predicated store instruction comprises a compiler hint field, and the prefetch logic only initiates the memory operation when indicated by the compiler hint field.
9 . The block-based processor core of claim 1 , wherein the initiated memory operation is prioritized behind non-prefetch requests to the memory hierarchy.
10 . The block-based processor core of claim 1 , further comprising:
wake-up and select logic configured to determine when the first value associated with the predicated store instruction is ready and to initiate the prefetch logic after the first value is ready.
11 . A method of executing a program on a processor comprising a block-based processor core, the method comprising:
receiving an instruction block comprising a plurality of instructions; determining that an instruction of the plurality of instructions is a predicated store instruction; and initiating a memory operation associated with a memory address targeted by the predicated store instruction before a predicate of the predicated store instruction is calculated.
12 . The method of claim 11 , further comprising:
calculating the memory address using a first value encoded in a field of the predicated store instruction and a second value generated by a register read or a different instruction targeting the predicated store instruction.
13 . The method of claim 11 , wherein initiating the memory operation comprises performing a cache coherence operation corresponding to a cache line including the memory address.
14 . The method of claim 11 , wherein initiating the memory operation comprises calculating the memory address comprises using a dedicated arithmetic unit.
15 . The method of claim 11 , wherein the memory operation is initiated only when indicated by a prefetch enable bit of the predicated store instruction.
16 . A method comprising:
receiving instructions of a program; grouping the instructions into a plurality of instruction blocks targeted for execution on a block-based processor; for a respective instruction block of the plurality of instruction blocks:
determine whether a store instruction is predicated;
classify a given predicated store instruction as a candidate for prefetching or not a candidate for prefetching; and
enable prefetching for the given predicated store instruction when it is classified as a candidate for prefetching;
emitting the plurality of instruction blocks for execution by the block-based processor; and storing the emitted plurality of instruction blocks in one or more computer-readable storage media or devices.
17 . The method of claim 16 , wherein classifying the given predicated store instruction is based only on static information about the program.
18 . The method of claim 17 , wherein classifying the given predicated store instruction is based on an instruction mix of the respective instruction block.
19 . The method of claim 16 , wherein classifying the given predicated store instruction is based on dynamic information about the program.
20 . One or more computer-readable storage media storing computer-readable instructions that when executed by a computer cause the computer to perform the method of claim 16 .Join the waitlist — get patent alerts
Track US2017083339A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.