US2017083339A1PendingUtilityA1

Prefetching associated with predicated store instructions

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Sep 19, 2015Filed: Mar 4, 2016Published: Mar 23, 2017
Est. expirySep 19, 2035(~9.1 yrs left)· nominal 20-yr term from priority
G06F 2212/602G06F 9/3891G06F 9/3013G06F 9/3822G06F 9/3836G06F 9/3824G06F 9/30076G06F 2212/62G06F 9/30189G06F 9/383G06F 9/30087G06F 2212/452G06F 9/3848G06F 9/3804G06F 9/268G06F 12/0811G06F 13/4221G06F 9/30101G06F 9/3016G06F 12/1009G06F 9/3838G06F 11/36G06F 9/3004G06F 9/3009G06F 9/466G06F 9/30105G06F 9/32G06F 9/3828G06F 9/30047G06F 12/0875G06F 9/30072G06F 9/355G06F 9/30021G06F 15/7867G06F 9/30007G06F 15/80G06F 9/30167G06F 11/3648G06F 2212/604G06F 9/30145G06F 9/3853G06F 9/30098G06F 12/0806G06F 11/3656G06F 9/3802G06F 9/3557G06F 15/8007G06F 9/345G06F 12/0862G06F 9/3867G06F 9/35G06F 9/528G06F 9/30043G06F 9/321G06F 9/3842G06F 9/3851G06F 9/30058G06F 9/3005G06F 9/30038G06F 9/3858G06F 9/3854G06F 9/30138Y02D10/00G06F 9/3856G06F 9/38585
39
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Technology related to prefetching data associated with predicated stores of programs in block-based processor architectures is disclosed. In one example of the disclosed technology, a processor includes a block-based processor core for executing an instruction block comprising a plurality of instructions. The block-based processor core includes decode logic and prefetch logic. The decode logic is configured to detect a predicated store instruction of the instruction block. The prefetch logic is configured to calculate a target address of the predicated store instruction and initiate a memory operation associated with the calculated target address before a predicate of the predicated store instruction is calculated.

Claims

exact text as granted — not AI-modified
We claim: 
     
         1 . A processor comprising a block-based processor core for executing an instruction block comprising an instruction header and a plurality of instructions, the block-based processor core comprising:
 decode logic configured to detect a predicated store instruction of the instruction block; and   prefetch logic configured to:
 receive a first value associated with the predicated store instruction; 
 calculate a target address of the predicated store instruction using the received first value; and 
 initiate a memory operation associated with the calculated target address before a predicate of the predicated store instruction is calculated. 
   
     
     
         2 . The block-based processor core of  claim 1 , wherein the memory operation includes issuing a prefetch request to a memory hierarchy of the processor to prefetch a cache line spanning the calculated target address. 
     
     
         3 . The block-based processor core of  claim 1 , wherein the memory operation includes fetching coherence permissions for a memory line including data at the calculated target address. 
     
     
         4 . The block-based processor core of  claim 1 , wherein the memory operation includes determining whether an inter-thread conflict exists for a memory line spanning the calculated target address. 
     
     
         5 . The block-based processor core of  claim 1 , wherein the target address is calculated using a dedicated arithmetic unit of the prefetch logic. 
     
     
         6 . The block-based processor core of  claim 1 , wherein the first value is generated by another instruction of the instruction block that targets the predicated store instruction. 
     
     
         7 . The block-based processor core of  claim 1 , wherein calculating the target address comprises performing the target address calculation during an open instruction issue slot and using an arithmetic unit of instruction execution logic. 
     
     
         8 . The block-based processor core of  claim 1 , wherein the predicated store instruction comprises a compiler hint field, and the prefetch logic only initiates the memory operation when indicated by the compiler hint field. 
     
     
         9 . The block-based processor core of  claim 1 , wherein the initiated memory operation is prioritized behind non-prefetch requests to the memory hierarchy. 
     
     
         10 . The block-based processor core of  claim 1 , further comprising:
 wake-up and select logic configured to determine when the first value associated with the predicated store instruction is ready and to initiate the prefetch logic after the first value is ready.   
     
     
         11 . A method of executing a program on a processor comprising a block-based processor core, the method comprising:
 receiving an instruction block comprising a plurality of instructions;   determining that an instruction of the plurality of instructions is a predicated store instruction; and   initiating a memory operation associated with a memory address targeted by the predicated store instruction before a predicate of the predicated store instruction is calculated.   
     
     
         12 . The method of  claim 11 , further comprising:
 calculating the memory address using a first value encoded in a field of the predicated store instruction and a second value generated by a register read or a different instruction targeting the predicated store instruction.   
     
     
         13 . The method of  claim 11 , wherein initiating the memory operation comprises performing a cache coherence operation corresponding to a cache line including the memory address. 
     
     
         14 . The method of  claim 11 , wherein initiating the memory operation comprises calculating the memory address comprises using a dedicated arithmetic unit. 
     
     
         15 . The method of  claim 11 , wherein the memory operation is initiated only when indicated by a prefetch enable bit of the predicated store instruction. 
     
     
         16 . A method comprising:
 receiving instructions of a program;   grouping the instructions into a plurality of instruction blocks targeted for execution on a block-based processor;   for a respective instruction block of the plurality of instruction blocks:
 determine whether a store instruction is predicated; 
 classify a given predicated store instruction as a candidate for prefetching or not a candidate for prefetching; and 
 enable prefetching for the given predicated store instruction when it is classified as a candidate for prefetching; 
   emitting the plurality of instruction blocks for execution by the block-based processor; and   storing the emitted plurality of instruction blocks in one or more computer-readable storage media or devices.   
     
     
         17 . The method of  claim 16 , wherein classifying the given predicated store instruction is based only on static information about the program. 
     
     
         18 . The method of  claim 17 , wherein classifying the given predicated store instruction is based on an instruction mix of the respective instruction block. 
     
     
         19 . The method of  claim 16 , wherein classifying the given predicated store instruction is based on dynamic information about the program. 
     
     
         20 . One or more computer-readable storage media storing computer-readable instructions that when executed by a computer cause the computer to perform the method of  claim 16 .

Join the waitlist — get patent alerts

Track US2017083339A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.