US2016011642A1PendingUtilityA1

Power and throughput optimization of an unbalanced instruction pipeline

Assignee: TEXAS INSTRUMENTS INCPriority: Apr 20, 2010Filed: Sep 21, 2015Published: Jan 14, 2016
Est. expiryApr 20, 2030(~3.7 yrs left)· nominal 20-yr term from priority
G06F 11/3466G06F 9/3869G06F 1/324Y02D10/00G06F 2201/885G06F 11/3409G06F 2201/88
36
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method includes determining a rate of resource occupancy of a constituent stage of an unbalanced instruction pipeline implemented in a processor through profiling an instruction code. The method also includes performing data processing at a maximum throughput at an optimum clock frequency based on the rate of resource occupancy.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 determining a rate of resource occupancy of a constituent stage of an unbalanced instruction pipeline implemented in a processor through profiling an instruction code; and   performing data processing associated with the unbalanced instruction pipeline at a maximum throughput and at an optimum clock frequency based on the rate of resource occupancy.   
     
     
         2 . The method of  claim 1 , wherein performing the data processing includes stalling processing associated with at least one of the constituent stage of the unbalanced instruction pipeline and a previous stage, for at least a number of clock cycles corresponding to a delay time associated with the processing, through the constituent stage by gating a clock input to the at least one of the constituent stage and the previous stage. 
     
     
         3 . The method of  claim 1 , further comprising:
 determining a time interval within a processing time associated with the constituent stage of the unbalanced instruction pipeline based on a change in a processing scenario associated with processing;   dynamically determining the rate of resource occupancy of the constituent stage periodically with a time period equal to the determined time interval; and   obtaining, at every time interval, the clock frequency associated with the rate of resource occupancy of the constituent stage for performing the data processing associated with the unbalanced instruction pipeline.   
     
     
         4 . The method of  claim 3 , wherein the clock frequency associated with the data processing is higher than a frequency corresponding to the higher delay time associated with the constituent stage. 
     
     
         5 . The method of  claim 2 , further comprising obtaining a number of stall cycles associated with stalling in at least one of the constituent stage of the unbalanced instruction pipeline and a previous stage thereof for at least the number of stall cycles, wherein the number of stall cycles corresponds to a delay time associated with the processing through the constituent stage. 
     
     
         6 . The method of  claim 5 , wherein determining the rate of resource occupancy of the constituent stage of the unbalanced instruction pipeline includes:
 inputting a control signal associated with a decoded instruction associated with the processing through the constituent stage to a counter associated therewith;   determining the rate of resource occupancy of the constituent stage through the counter; and   maintaining a Look Up Table (LUT) associated with the counter to map the determined rate of resource occupancy and at least one of the clock frequency and the number of stall cycles associated therewith.   
     
     
         7 . The method of  claim 6 , further comprising:
 updating hardware associated with the processing through the constituent stage with the at least one of the clock frequency and the number of stall cycles determined through the LUT when the at least one of the clock frequency and the number of stall cycles varies from a value thereof during a previous time interval; and   resetting the counter at the end of the time interval.   
     
     
         8 . The method of  claim 6 , comprising implementing the LUT through a multiplexer having the rate of resource occupancy as an input and a select line. 
     
     
         9 . A method comprising:
 determining a time interval within a processing time associated with a constituent stage of an unbalanced instruction pipeline implemented in a processor based on a change in a processing scenario associated with data processing;   dynamically determining a rate of resource occupancy of the constituent stage periodically with a time period equal to the time interval through profiling an instruction code;   periodically obtaining a clock frequency associated with the rate of resource occupancy of the constituent stage, the clock frequency corresponding to an optimized at least one of a power consumption and a throughput associated with the unbalanced instruction pipeline; and   performing the data processing at the periodically obtained clock frequency.   
     
     
         10 . The method of  claim 9 , further comprising obtaining a number of stall cycles associated with stalling processing in at least one of the constituent stage of the unbalanced instruction pipeline and a previous stage thereof for at least the number of stall cycles,
 wherein the number of stall cycles corresponds to a delay time associated with the processing through the constituent stage.   
     
     
         11 . The method of  claim 9 , wherein dynamically determining the rate of resource occupancy of the constituent stage includes:
 inputting a control signal associated with a decoded instruction associated with the processing through the constituent stage to a counter associated therewith;   determining the rate of resource occupancy of the constituent stage through the counter; and   maintaining a Look Up Table (LUT) associated with the counter to map the determined rate of resource occupancy and at least one of the clock frequency and the number of stall cycles associated therewith.   
     
     
         12 . The method of  claim 11 , further comprising:
 updating hardware associated with the processing through the constituent stage with the at least one of the clock frequency and the number of stall cycles determined through the LUT when the at least one of the clock frequency and the number of stall cycles varies from a value thereof during a previous time interval; and   resetting the counter at the end of the time interval.   
     
     
         13 . The method of  claim 11 , comprising implementing the LUT through a multiplexer having the rate of resource occupancy as an input and a select line thereof. 
     
     
         14 . A computing system comprising:
 a processor having an unbalanced instruction pipeline;   a memory configured to store an instruction code associated with processing through the unbalanced instruction pipeline; and   a determination module configured to determine a rate of resource occupancy of a constituent stage of the unbalanced instruction pipeline through profiling the instruction code associated with processing through the unbalanced instruction pipeline, the processor being configured to perform data processing at a maximum throughput at an optimum clock frequency based on the rate of resource occupancy.   
     
     
         15 . The computing system of  claim 14 , further comprising a pipeline control unit configured to control a clock generation circuit associated with the constituent stage of the unbalanced instruction pipeline. 
     
     
         16 . The computing system of  claim 16 , wherein the pipeline control unit further comprises a Look Up Table (LUT) implemented therein configured to map the rate of resource occupancy of the constituent stage determined through the determination module to at least one of the clock frequency and a number of stall cycles,
 wherein the number of stall cycles is associated with stalling processing in at least one of the constituent stage of the unbalanced instruction pipeline and a previous stage thereof for at least the number of stall cycles, and   wherein the number of stall cycles corresponds to a delay time associated with the processing through the constituent stage.

Join the waitlist — get patent alerts

Track US2016011642A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.