Adaptive Cache Partitioning
Abstract
Described apparatuses and methods partition a cache memory based, at least in part, on a metric indicative of prefetch performance. The amount of cache memory allocated for metadata related to prefetch operations versus cache storage can be adjusted based on operating conditions. Thus, the cache memory can be partitioned into a first portion allocated for metadata pertaining to an address space (prefetch metadata) and a second portion allocated for data associated with the address space (cache data). The amount of cache memory allocated to the first portion can be increased under workloads that are suitable for prefetching and decreased otherwise. The first portion may include one or more cache units, cache lines, cache ways, cache sets, or other resources of the cache memory.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A method comprising:
implementing a partitioning scheme to partition a cache memory into a first portion and a second portion; servicing multiple requests relating to an address space, including:
maintaining metadata pertaining to the address space within the first portion of the cache memory, and
loading data associated with addresses of the address space into the second portion of the cache memory; and
modifying the partitioning scheme to adapt a size of the first portion based, at least in part, on a metric quantifying performance of the loading.
22 . The method of claim 21 , further comprising:
implementing a metadata mapping scheme to access cache units allocated to the first portion of the cache memory; and implementing an address mapping scheme to map addresses of the address space to cache units allocated to the second portion of the cache memory.
23 . The method of claim 21 , further comprising:
loading data into the second portion of the cache memory in response to cache misses, including requests pertaining to addresses that are not available within the cache memory.
24 . The method of claim 21 , further comprising:
prefetching data into the second portion of the cache memory based, at least in part, on the metric quantifying performance of the loading.
25 . The method of claim 24 , further comprising:
utilizing the metadata maintained within the first portion of the cache memory to predict addresses of upcoming requests; and prefetching data corresponding to the predicted addresses into the second portion of the cache memory before requests pertaining to the predicted addresses are received at the cache memory.
26 . The method of claim 21 , further comprising:
modifying the partitioning scheme to adapt a size of the second portion based, at least in part, on the metric quantifying performance of the loading.
27 . The method of claim 26 , further comprising:
modifying the partitioning scheme to adapt the size of the first portion and the size of the second portion based, at least in part, on a metric quantifying prefetch performance of the loading.
28 . The method of claim 27 , further comprising:
monitoring the metric quantifying prefetch performance that pertains to data prefetched into the second portion of the cache memory; and comparing the metric quantifying prefetch performance to one or more thresholds.
29 . The method of claim 21 , further comprising:
increasing the size of the first portion of the cache memory allocated for the metadata and decreasing a size of the second portion of the cache memory allocated for the data responsive to the metric quantifying performance exceeding at least one threshold.
30 . The method of claim 21 , further comprising:
increasing the size of the first portion and decreasing a size of the second portion, including compacting data stored within the second portion of the cache memory; and decreasing the size of the first portion and increasing the size of the second portion, including compacting metadata stored within the first portion of the cache memory.
31 . An apparatus comprising:
a memory array configured as a cache memory; and logic coupled to the memory array, the logic configured to:
implement a partitioning scheme to partition the cache memory into a first portion and a second portion;
service multiple requests relating to an address space, including:
maintaining metadata pertaining to the address space within the first portion of the cache memory, and
loading data associated with addresses of the address space into the second portion of the cache memory; and
modify the partitioning scheme to adapt a size of the first portion based, at least in part, on a metric quantifying performance of the loading.
32 . The apparatus of claim 31 , wherein the logic is further configured to:
maintain within the first portion of the cache memory one or more of an address sequence, address history, index table, delta sequence, stride pattern, correlation pattern, feature vector, machine-learned (ML) feature, ML feature vector, ML model, or ML modeling data.
33 . The apparatus of claim 31 , wherein:
the metric quantifying performance of the loading comprises a prefetch metric; and the logic is further configured to monitor one or more of a prefetch hit rate, quantity of useful prefetches, quantity of bad prefetches, or ratio of useful prefetches to bad prefetches.
34 . The apparatus of claim 31 , wherein the logic is further configured to:
increase the size of the first portion of the cache memory for the metadata pertaining to the address space in response to the metric exceeding at least one threshold; and decrease a size of the second portion of the cache memory for the data associated with addresses of the address space in response to the metric exceeding the at least one threshold.
35 . The apparatus of claim 34 , wherein the logic is further configured to:
decrease the size of the first portion of the cache memory for the metadata pertaining to the address space in response to the metric being below the at least one threshold; and increase the size of the second portion of the cache memory for the data associated with addresses of the address space in response to the metric being below the at least one threshold.
36 . The apparatus of claim 31 , wherein the logic is further configured to:
allocate a quantity of ways of the cache memory for the metadata pertaining to the address space; and modify the quantity of ways of the cache memory allocated for the metadata pertaining to the address space based, at least in part, on the metric quantifying performance of the loading.
37 . The apparatus of claim 36 , wherein the logic is further configured to:
divide ways of a set of the cache memory into a first group allocated for the metadata pertaining to the address space and a second group allocated for the data associated with addresses of the address space; and move data from a way within the first group to a way within the second group.
38 . The apparatus of claim 31 , wherein the logic is further configured to:
allocate a quantity of sets of the cache memory for the metadata pertaining to the address space; and modify the quantity of sets of the cache memory allocated for the metadata pertaining to the address space based, at least in part, on the metric quantifying performance of the loading.
39 . The apparatus of claim 31 , wherein the logic is further configured to:
responsive to allocating a group of cache units from the second portion to the first portion,
evicting data from cache units of the group of cache units,
disabling cache tags associated with the cache units of the group of cache units,
evicting data from a selected cache unit, the selected cache unit to remain allocated to the second portion, and
moving to the selected cache unit data stored within a cache unit of the group of cache units being allocated from the second portion to the first portion.
40 . An apparatus comprising:
a memory array configured as a cache memory; means for implementing a partitioning scheme to partition the cache memory into a first portion and a second portion; means for maintaining metadata pertaining to an address space within the first portion of the cache memory; means for loading data associated with addresses of the address space into the second portion of the cache memory; and means for modifying the partitioning scheme to adapt a size of the first portion based, at least in part, on a metric quantifying performance of the loading of the data associated with addresses of the address space.Join the waitlist — get patent alerts
Track US2023169011A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.