Intermediate Representation Controller Circuit for Selecting Hardware Compute Units to Process Microcode According to Identified Intermediate Representation Primitives
Abstract
An intermediate representation (IR) controller is described that, for a given intermediate representation (IR) primitive, selects a hardware compute unit of a plurality of hardware compute units. In a non-limiting example, the IR controller receives an input that specifies an IR primitive, a device mask indicating a type of hardware circuitry to be used to process the primitive, and a goal vector specifying a goal in the processing of the primitive. The IR controller also collects data describing power consumption by respective hardware compute units and completion times for processing respective IR primitives. This data is maintained as implementation profiles that describe operation of respective hardware compute units in processing respective IR primitives, e.g., as histograms. The implementation profiles are then leveraged by the IR controller to select hardware compute units for execution of subsequent IR primitives.
Claims
exact text as granted — not AI-modified1 . A method comprising:
receiving, by a controller circuit, from a neural network, an input identifying a goal vector and a specific intermediate representation (IR) primitive of a plurality of IR primitives, the goal vector defining a goal in how the specific IR primitive is to be processed; identifying, by the controller circuit, at least one implementation profile from a plurality of implementation profiles based on the input, the plurality of implementation profiles describing operation of a plurality of microcode implementations in processing respective ones of the plurality of IR primitives; selecting, by the controller circuit, a microcode implementation from the plurality of microcode implementations based on the at least one implementation profile and the goal; and invoking, by the controller circuit, processing of microcode corresponding to the specific IR primitive by the selected microcode implementation.
2 . The method of claim 1 , wherein the selecting of the microcode implementation from the plurality of microcode implementations is based at least in part on the goal in how the specific IR primitive is to be processed.
3 . The method of claim 2 , wherein the goal in how the specific IR primitive is to be processed specifies whether performance or power efficiency is to be given a relatively higher priority when implementing the specific IR primitive.
4 . The method of claim 1 , wherein the input further identifies a device mask specifying a type of hardware circuitry to be used to process the specific IR primitive and the identifying or the selecting is based at least in part on the type.
5 . (canceled)
6 . The method of claim 1 , further comprising detecting, by the controller circuit, operating conditions of hardware compute units corresponding to the plurality of microcode implementations and wherein the selecting is based at least in part on the detected operating conditions.
7 . The method of claim 6 , wherein the hardware compute units are implemented by a central processing unit, parallel processing unit, floating point grid array, or tensor processing unit.
8 . The method of claim 1 , further comprising:
receiving, by the controller circuit, data describing operation of the selected microcode implementation in processing the microcode; and updating, by the controller circuit, the at least one implementation profile based on the data.
9 . The method of claim 1 , further comprising updating, by the controller circuit, one or more of the plurality of implementation profiles offline.
10 . The method of claim 1 , wherein the plurality of implementation profiles are maintained as part of writeable microcode in a writeable control store.
11 . An intermediate representation (IR) controller circuit comprising:
an input module circuit configured to receive an input from a neural network, the input identifying a goal vector and an IR primitive, and the goal vector defining a goal in how the IR primitive is to be processed; a profiler manager module circuit configured to collect data in a writeable control store as a plurality of implementation profiles, the plurality of implementation profiles describing operation of a plurality of hardware compute units in processing, respectively, a plurality of microcode implementations; and an actuator module circuit configured to select, based on the goal, a specific hardware compute unit of the plurality of hardware compute units to process microcode corresponding to the IR primitive.
12 . (canceled)
13 . The IR controller circuit of claim 11 , wherein the plurality of implementation profiles describe the operation using histograms.
14 . The IR controller circuit of claim 13 , wherein the histograms describe power consumption or performance.
15 . The IR controller circuit of claim 11 , wherein the actuator module circuit is configured to select the specific hardware compute unit based on operating conditions detected for the plurality of hardware compute units.
16 . The IR controller circuit of claim 11 , wherein the goal in how the IR primitive is to be processed specifies whether performance or power efficiency is to be given a relatively higher priority when implementing the IR primitive.
17 . The IR controller circuit of claim 11 , wherein the input further identifies a device mask specifying a type of hardware circuitry to be used to process the IR primitive and the actuator module circuit is configured to select the hardware compute unit from the plurality of hardware compute units based at least in part on the type.
18 . A method comprising:
generating a plurality of implementation profiles by an intermediate representation (IR) controller circuit based on data collected from and describing operation of a plurality of hardware compute units in processing microcode corresponding to an IR primitive; forming an additional implementation profile by the IR controller circuit based on data collected from an additional hardware compute unit made available by communicatively coupling the additional hardware compute unit to the IR controller circuit; receiving an input at the IR controller circuit to cause processing of the IR primitive; determining by the IR controller circuit which hardware compute unit of the plurality of hardware compute units, including the additional hardware compute unit, is to be used to process the IR primitive based on the plurality of implementation profiles and the additional implementation profile; and invoking processing of microcode corresponding to the IR primitive at the hardware compute unit by the IR controller circuit.
19 . The method of claim 18 , wherein the forming, the receiving, the determining, and the invoking are performed in real time.
20 . The method of claim 18 , wherein the generating is performed offline while at least a portion of the plurality of hardware compute units are idled.
21 . The method of claim 18 , wherein receiving the input comprises receiving the input from machine learning software.
22 . The method of claim 21 , wherein receiving the input comprises receiving the input from a neural network.Join the waitlist — get patent alerts
Track US2024004645A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.