US2020233664A1PendingUtilityA1

Efficient range-based memory writeback to improve host to device communication for optimal power and performance

Assignee: INTEL CORPPriority: Mar 31, 2017Filed: Dec 17, 2019Published: Jul 23, 2020
Est. expiryMar 31, 2037(~10.7 yrs left)· nominal 20-yr term from priority
G06F 9/30047Y02D10/00G06F 12/0891G06F 12/128G06F 2212/62G06F 12/0875G06F 9/3834G06F 12/0808G06F 13/28G06F 2212/1016G06F 2212/452G06F 9/3016G06F 12/0811G06F 12/0804G06F 9/30087G06F 12/084G06F 12/0842G06F 9/3004Y02D10/14G06F 9/3858
59
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Method and apparatus for efficient range-based memory writeback is described herein. One embodiment of an apparatus includes a system memory, a plurality of hardware processor cores each of which includes a first cache, a decoder circuitry to decode an instruction having fields for a first memory address and a range indicator, and an execution circuitry to execute the decoded instruction. Together, the first memory address and the range indicator define a contiguous region in the system memory that includes one or more cache lines. An execution of the decoded instruction causes any instances of the one or more cache lines in the first cache to be invalidated. Additionally, any invalidated instances of the one or more cache lines that are dirty are to be stored to system memory.

Claims

exact text as granted — not AI-modified
1 .- 22 . (canceled) 
     
     
         23 . An apparatus comprising:
 a shared cache shared by a first processing unit and a second processing unit; and   a private cache private to the first processing unit to store a plurality of cache lines;   wherein a subset of the plurality of cache lines is copied from the private cache to the shared cache responsive to an execution of an instruction by the first processing unit,   wherein the subset of cache lines comprises every dirty cache line within the plurality of cache lines; and   wherein each of the plurality of cache lines comprises a memory address within a contiguous address range defined by a starting memory address and a range indicator specified by the instruction.   
     
     
         24 . The apparatus of  claim 23 , wherein a cache coherence state of each of the plurality of cache lines in the private cache is set to an Invalidate state responsive to the execution of the instruction by the first processing unit. 
     
     
         25 . The apparatus of  claim 23 , wherein a cache coherence state of each of the subset of cache line is set to a Shared state responsive to the execution of the instruction by the first processing unit. 
     
     
         26 . The apparatus of  claim 23 , wherein the range indicator comprises an ending memory address and the contiguous address range is to span from the starting memory address to the ending memory address. 
     
     
         27 . The apparatus of  claim 23 , wherein the range indicator comprises a number of incrementally-addressed cache lines to be included in the contiguous address range starting with a first cache line having a memory address equal to, or larger than, the starting memory address. 
     
     
         28 . The apparatus of  claim 23 , wherein the starting memory address is a linear memory address. 
     
     
         29 . A method comprising:
 storing a plurality of cache lines in a private cache private to a first processing unit;   copying a subset of the plurality of cache lines from the private cache to a shared cache responsive to an execution of an instruction by the first processing unit, the shared cache shared by the first processing unit and a second processing unit;   wherein the subset of cache lines comprises every dirty cache line within the plurality of cache lines; and   wherein each of the plurality of cache lines comprises a memory address within a contiguous address range defined by a starting memory address and a range indicator specified by the instruction.   
     
     
         30 . The method of  claim 29 , further comprising setting a cache coherence state of each of the plurality of cache lines in the private cache to an Invalidate state responsive to the execution of the instruction by the first processing unit. 
     
     
         31 . The method of  claim 29 , further comprising setting a cache coherence state of each of the subset of cache lines in the private cache to a shared state responsive to the execution of the instruction by the first processing unit. 
     
     
         32 . The method of  claim 29 , wherein the range indicator comprises an ending memory address and the contiguous address range is to span from the starting memory address to the ending memory address. 
     
     
         33 . The method of  claim 29 , wherein the range indicator comprises a number of incrementally-addressed cache lines to be included in the contiguous address range starting with a first cache line having a memory address equal to, or larger than, the starting memory address. 
     
     
         34 . The method of  claim 29 , wherein the starting memory address is a linear memory address. 
     
     
         35 . A non-transitory machine-readable medium having program code stored thereon which, when executed by a machine, causes the machine to perform operations of:
 storing a plurality of cache lines in a private cache private to a first processing unit;   copying a subset of the plurality of cache lines from the private cache to a shared cache responsive to an execution of an instruction by the first processing unit, the shared cache shared by the first processing unit and a second processing unit;   wherein the subset of cache lines comprises every dirty cache line within the plurality of cache lines; and   wherein each of the plurality of cache lines comprises a memory address within a contiguous address range defined by a starting memory address and a range indicator specified by the instruction.   
     
     
         36 . The non-transitory machine-readable medium of  claim 35 , wherein the operations further comprise setting a cache coherence state of each of the plurality of cache lines in the private cache to an Invalidate state responsive to the execution of the instruction by the first processing unit. 
     
     
         37 . The non-transitory machine-readable medium of  claim 35 , wherein the operations further comprise setting a cache coherence state of each of the subset of cache lines in the private cache to a shared state responsive to the execution of the instruction by the first processing unit. 
     
     
         38 . The non-transitory machine-readable medium of  claim 35 , wherein the range indicator comprises an ending memory address and the contiguous address range is to span from the starting memory address to the ending memory address. 
     
     
         39 . The non-transitory machine-readable medium of  claim 35 , wherein the range indicator comprises a number of incrementally-addressed cache lines to be included in the contiguous address range starting with a first cache line having a memory address equal to, or larger than, the starting memory address. 
     
     
         40 . The non-transitory machine-readable medium of  claim 35 , wherein the starting memory address is a linear memory address.

Join the waitlist — get patent alerts

Track US2020233664A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.