Method and apparatus for minimally intrusive instruction pointer-aware processing resource activity profiling
Abstract
Systems and methods for minimally intrusive instruction pointer-aware processing resource activity profiling are disclosed. In one embodiment, a graphics processor includes a grouping of processing resources and control logic that is associated with the grouping of processing resources. The control logic is configured to sample a state of at least one processing resource of the grouping of processing resources and to determine activity data from the state with the activity data including at least one of stalls and reason counts for stalling activity, instruction types, pipeline utilization, thread utilization, and shader activity.
Claims
exact text as granted — not AI-modified1 . A graphics processor, comprising:
a grouping of processing resources; and control logic that is associated with the grouping of processing resources, the control logic is configured to sample a state of at least one processing resource of the grouping of processing resources and to determine activity data from the state with the activity data including reason counts for counting occurrences for stalling activity, or activity counts for instruction types, pipeline utilization, thread utilization, or shader activity.
2 . The graphics processor of claim 1 , further comprising:
a cache unit that is associated with the grouping of processing resources, the cache unit to receive an instruction pointer address and the activity data including a stall reason for each state of processing resources that are associated with the cache unit.
3 . The graphics processor of claim 2 , wherein each sampling of a state is scheduled for a chosen clock cycle and is minimally intrusive.
4 . The graphics processor of claim 1 , wherein the control logic is configured to store a state when threads are allocated on a processing resource with no instruction being executed for a chosen cycle that is sampled.
5 . The graphics processor of claim 4 , wherein the control logic is configured to discard a state for a chosen cycle that is sampled if the processing resource is idle or executing an instruction.
6 . The graphics processor of claim 1 , wherein the control logic is configured to interleave samplings of states of processing resources among the grouping of processing resources and other groupings of processing resources, to resolve the states into one of a number of supported stall reasons, and to prioritize the supported stall reasons based on a priority level of the stall reasons.
7 . The graphics processor of claim 1 , wherein the supported stalls and reason counts for stalling activity comprise a synch stall field for a stall or delay between threads to reach a common point, an instruction fetch field for an instruction fetch from memory that is stalled, a scoreboard field for a stall based on a data dependency, a send stall field for a send bus bandwidth limit for an processing resource, a pipe stall field for a stall within a pipeline, and an internal stall field for a stall caused from a memory bank collision.
8 . A cache structure, comprising:
logic to perform operations of the cache structure; and memory coupled to the logic, the memory to store instruction pointer addresses and associated data fields to indicate activity data from sampling of processing resources, wherein the logic is configured to receive an instruction pointer address and activity data for a state of processing resources that are associated with the cache structure.
9 . The cache structure of claim 8 , wherein the logic is configured to perform an instruction pointer address lookup within the cache structure.
10 . The cache structure of claim 9 , wherein the logic is configured to build an entry for a new cache line when the instruction pointer lookup misses, to store the instruction pointer address and the activity data in the new cache line, to initialize the identified activity including a stall reason to a count while all other reason counts are initialized to a different count.
11 . The cache structure of claim 10 , wherein the logic is configured to determine if all available lines of the cache structure are occupied and to perform a capacity-eviction to evict an existing line if all available lines of the cache structure are occupied.
12 . The cache structure of claim 9 , wherein the logic is configured to determine a hit for instruction pointer address lookup, to perform a read operation of a cache line for the instruction pointer address, to perform a modify operation to increment a count of the identified activity, and to perform a write operation for the cache line.
13 . The cache structure of claim 9 , wherein the logic is configured for a maximum value eviction when a given cache line has an activity count that reaches a maximum representable value and performs the maximum value eviction by evicting the instruction pointer address and its corresponding data to a circular buffer in main memory.
14 . A method for minimally intrusive profiling of a graphics processing unit (GPU), comprising:
receiving, with a cache unit, an instruction pointer address and activity data for each state of processing resources that are associated with the cache unit; and performing an instruction pointer address lookup within the cache unit for the received instruction pointer address and associated activity data.
15 . The method of claim 14 , further comprising:
building an entry for a new cache line when the instruction pointer lookup misses.
16 . The method of claim 15 , further comprising:
storing the instruction pointer address and the activity data in the new cache line; and initializing the identified activity including a stall reason to 1 while all other reason counts are initialized to 0.
17 . The method of claim 16 , further comprising:
determining if all available lines of the cache structure are occupied and to perform a capacity-eviction to evict an existing line if all available lines of the cache structure are occupied.
18 . The method of claim 15 , further comprising:
determining a hit for instruction pointer address lookup.
19 . The method of claim 18 , further comprising:
performing a read operation of a cache line for the instruction pointer address; performing a modify operation to increment a count of the identified activity for the instruction pointer address; and performing a write operation for the cache line.
20 . The method of claim 19 , further comprising:
performing a maximum value eviction when a given cache line has an activity count that reaches a maximum representable value, wherein performing the maximum value eviction comprises evicting the instruction pointer address and its corresponding data to a circular buffer in main memory.
21 . (canceled)Join the waitlist — get patent alerts
Track US2022156068A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.