Efficient task allocation
Abstract
A method and processor comprising a command processing unit to receive, from a host processor, a sequence of commands to be executed; and generate based on the sequence of commands a plurality of tasks. The processor also comprises a plurality of compute units each having a first processing module for executing tasks of a first task type, a second processing module for executing tasks of a second task type, different from the first task type, and a local cache shared by at least the first processing module and the second processing module. The command processing unit issues the plurality of tasks to at least one of the plurality of compute units, and wherein at least one of the plurality of compute units is to process at least one of the plurality of tasks.
Claims
exact text as granted — not AI-modified1 . A processor comprising:
a command processing unit to:
receive, from a host processor, a sequence of commands to be executed; and
generate based on the sequence of commands a plurality of tasks; and
a plurality of compute units, wherein at least one of the plurality of compute units comprises:
a first processing module for executing tasks of a first task type generated by the command processing unit;
a second processing module for executing tasks of a second task type, different from the first task type, generated by the command processing unit;
a local cache shared by at least the first processing module and the second processing module;
wherein the command processing unit is to issue the plurality of tasks to at least one of the plurality of compute units, and wherein at least one of the plurality of compute units is to process at least one of the plurality of tasks.
2 . The processor of claim 1 , wherein the command processing unit is to issue tasks of the first task type to the first processing module of a given compute unit and to issue tasks of the second task type to the second processing module of the plurality of a given compute unit.
3 . The processor of claim 1 , wherein the first task type is a task for undertaking at least a portion of a graphics processing operation forming one of a set of pre-defined graphics processing operations which collectively enable the implementation of a graphics processing pipeline, and wherein the second task type is a task for undertaking at least a portion of a neural processing operation.
4 . The processor of claim 3 , wherein the graphics processing operation comprises at least one of:
a graphics compute shader task; a vertex shader task; a fragment shader task; a tessellation task; and a geometry shader task.
5 . The processor of claim 1 , wherein each compute unit is a shader core in a graphics processing unit.
6 . The processor of claim 1 , wherein the first processing module is a graphics processing module and wherein the second processing module is a neural processing module.
7 . The processor of claim 1 , wherein the command processing unit further comprises at least one dependency tracker to track dependencies between commands in the sequence of commands; and wherein the command processing unit is to use the at least one dependency tracker to wait for completion of processing of a given task of a first command in the sequence of commands before issuing an associated task of a second command in the sequence of commands for processing, where the associated task is dependent on the given task.
8 . The processor of claim 7 , wherein an output of the given task is stored in the local cache.
9 . The processor of claim 7 , wherein each command in the sequence of commands has metadata, wherein the metadata comprises indications of at least a number of tasks in the command, and task types associated with each of the tasks.
10 . The processor of claim 9 , wherein the command processing unit allocates each command in the sequence of commands, a command identifier, and the dependency tracker tracks dependencies between commands in the sequence of commands based on the command identifier.
11 . The processor of claim 10 , wherein when the given task of the first command is dependent on the associated task of the second command, the command processing unit allocates the given task and the associated task a same task identifier.
12 . The processor of claim 11 , wherein tasks of each of the commands that have been allocated the same task identifier are executed on the same compute unit of the plurality of compute units.
13 . The processor of claim 10 , wherein a task allocated a first task identifier is executed on a first compute unit of the plurality of compute units and a task allocated a second, different, task identifier is executed on a second compute unit of the plurality of compute units.
14 . The processor of claim 11 , wherein a task allocated a first task identifier, and of the first type, is executed on the first processing module of a given compute unit of the plurality of compute units, and a task allocated a second, different, task identifier, and of the second task type, is executed on the second processing module of the given compute unit of the plurality of compute units.
15 . The processor of claim 1 , wherein each of the plurality of compute units further comprise at least one queue of tasks, wherein the queue tasks comprise at least a part of the sequence of commands.
16 . The processor of claim 15 , wherein a given queue is associated with at least one task type.
17 . A method of allocating tasks associated with commands in a sequence of commands comprising:
receiving at a command processing unit, from a host processor, the sequence of commands to be executed; generating, at the command processing unit, based on the received sequence of commands a plurality of tasks; and issuing, by the command processing unit, each task to a compute unit of a plurality of compute units for execution, each compute unit comprising:
a first processing module for executing tasks of a first task type;
a second processing module for executing tasks of a second task type; and
a local cache shared by at least the first processing module and the second processing module;
wherein the command processing unit is to issue the plurality of tasks to at least one of the plurality of compute units, and wherein at least one of the plurality of compute units is to process at least one of the plurality of tasks.
18 . The method according to claim 17 , wherein the command processing unit waits for completion of processing of the tasks associated with the first command before issuing the tasks associated with the second command to the given compute unit, when the task associated with the second command is dependent on the task associated with the first command.
19 . The method according to claim 17 , wherein each command has associated metadata comprising indications of at least a number of tasks in the given command, and task types associated with each of the plurality of tasks.
20 . A non-transitory computer-readable storage medium comprising a set of computer-readable instructions stored thereon which, when executed by at least one processor are arranged to allocate tasks associated with commands in a sequence of commands wherein the instructions, when executed cause the at least one processor to:
receive at a command processing unit, from a host processor, the sequence of commands to be executed; generate, at the command processing unit, based on the received sequence of commands a plurality of tasks; and issue, by the command processing unit, each task to a compute unit of a plurality of compute units for execution, each compute unit comprising:
a first processing module for executing tasks of a first task type;
a second processing module for executing tasks of a second task type; and
a local cache shared by at least the first processing module and the second processing module;
wherein the command processing unit is to issue the plurality of tasks to at least one of the plurality of compute units, and wherein at least one of the plurality of compute units is to process at least one of the plurality of tasks.Join the waitlist — get patent alerts
Track US2024036919A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.