Method of reducing cache thrashing in a processing system and related processing system
Abstract
A method of reducing cache thrashing in a processing system is provided. M threads are issued to process a workload, and a memory access request associated with the M threads is transmitted to a first-level cache of the processing system. The memory access request is then transmitted to a second-level cache of the processing system in response to the first cache miss at the first-level cache. The memory access request is transmitted to a main memory of the processing system in response to the second cache miss at the second-level cache. The value of M is decreased when the relationship between the hit rates of the second-level cache and the first-level cache satisfies a predetermined criterion. A storage capacity and an access latency of the second-level cache are higher than those of the first-level cache.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of reducing cache thrashing in a processing system, comprising:
issuing M threads to process a workload; transmitting a memory access request associated with the M threads to a first-level cache of the processing system and determining whether a first cache hit or a first cache miss occurs at the first-level cache; transmitting the memory access request associated with the M threads to a second-level cache of the processing system in response to the first cache miss and determining whether a second cache hit or a second cache miss occurs at the second-level cache; transmitting the memory access request associated with the M threads to a main memory of the processing system in response to the second cache miss; determining whether a relationship between a first hit rate of the first-level cache and a second hit rate of the second-level cache satisfies a predetermined criterion; and decreasing a value of M when the predetermined criterion is satisfied, wherein:
a storage capacity of the second-level cache is higher than a storage capacity of the first-level cache;
an access latency of the second-level cache is higher than an access latency of the first-level cache; and
M is an integer larger than 1.
2 . The method of claim 1 , wherein:
the predetermined criterion is satisfied when the second hit rate is higher than the first hit rate.
3 . The method of claim 1 , wherein:
the predetermined criterion is satisfied when the second hit rate is higher than the first hit rate by more than a predetermined value.
4 . The method of claim 1 , wherein:
the predetermined criterion is satisfied when the second hit rate is higher than a predetermined positive value.
5 . The method of claim 1 , wherein:
a storage capacity of the main memory is higher than the storage capacity of the second-level cache; an access latency of the main memory is higher than the access latency of the second-level cache.
6 . The method of claim 1 , further comprising:
providing, by the first-level cache, data requested by the memory access request for completing an operation associated with the memory access request in response to the first cache hit.
7 . The method of claim 1 , further comprising:
providing, by the second-level cache, data requested by the memory access request for completing an operation associated with the memory access request in response to the second cache hit.
8 . The method of claim 1 , further comprising:
providing, by the main memory, data requested by the memory access request for completing an operation associated with the memory access request in response to the first cache miss and the second cache miss.
9 . The method of claim 1 , further comprising:
determining whether an adjustment made to the value of M has met a predetermined condition; and increasing the value of M when the predetermined condition is met.
10 . The method of claim 9 , wherein the predetermined condition is met when:
the value of M has been reduced more than K times; a difference between an original value of M and a current value of N exceeds a predetermined value; or a predetermined period of time has elapsed since a first decrease of the value of M.
11 . A processing system which reduces cache thrashing, comprising:
a plurality of processing cores; a first-level cache with a first storage capacity and a first access latency, and configured to:
receive a memory access request associated with M threads; and
determine whether a first cache hit or a first cache miss occurs at the first-level cache;
a second-level cache with a second storage capacity and a second access latency, and configured to:
receive the memory access request associated with the M threads from the first-level cache in response to a first cache miss at the first-level cache; and
determine whether a second cache hit or a second cache miss occurs at the second-level cache; and
a main memory configured to receive the memory access request associated with the M threads from the second-level cache in response to a second cache miss at the second-level cache; and a scheduler configured to:
issue the M threads to the plurality of processing cores for processing a workload;
transmit the memory access request associated with the M threads to the first-level cache;
determine whether a relationship between a first hit rate of the first-level cache and a second hit rate of the second-level cache satisfies a predetermined criterion; and
decrease a value of M when the predetermined criterion is satisfied, wherein:
the second storage capacity is higher than the first storage capacity;
the second access latency is higher than the first access latency; and
M is an integer larger than 1.
12 . The processing system of claim 11 , wherein:
the predetermined criterion is satisfied when the second hit rate is higher than the first hit rate.
13 . The processing system of claim 11 , wherein:
the predetermined criterion is satisfied when the second hit rate is higher than the first hit rate.
14 . The processing system of claim 11 , wherein:
the predetermined criterion is satisfied when the second hit rate is higher than a predetermined positive value.
15 . The processing system of claim 11 , wherein:
a third storage capacity of the main memory is higher than the second storage capacity; and a third access latency of the main memory is higher than the second access latency.
16 . The processing system of claim 11 , wherein the first-level cache is further configured to provide data requested by the memory access request to the plurality of processing cores in response to the first cache hit for completing the operation associated with the memory access request.
17 . The processing system of claim 11 , wherein the second-level cache is further configured to provide data requested by the memory access request to the plurality of processing cores in response to the second cache hit for completing the operation associated with the memory access request.
18 . The processing system of claim 11 , wherein the main memory is configured to provide data requested by the memory access request to the plurality of processing cores in response to the first cache miss and the second cache miss for completing the operation associated with the memory access request.
19 . The processing system of claim 11 , wherein the scheduler is further configured to:
determine whether an adjustment made to the value of M has met a predetermined condition; and increase the value of M when the predetermined condition is met.
20 . The processing system of claim 11 , wherein the plurality of processing cores, the scheduler and the first-level cache are implemented as a streaming multiprocessor (SM).Join the waitlist — get patent alerts
Track US2025252057A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.