Gpu networking using an integrated command processor
Abstract
Systems, apparatuses, and methods for generating network messages on a parallel processor are disclosed. A system includes at least a parallel processor, a general purpose processor, and a network interface unit. The parallel processor includes at least a plurality of compute units, a command processor, and a cache. A thread within a kernel executing on a compute unit of the parallel processor generates a network message and stores the network message and a corresponding indication in the cache. In response to detecting the indication of the network message in the cache, the command processor processes and conveys the network message to the network interface unit without involving the general purpose processor.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system comprising:
a first processor; and a second processor, wherein the second processor comprises a command processor, a plurality of compute units, and a cache; wherein the second processor is configured to:
generate a network message within a kernel executing on a compute unit;
store an indication of the network message in the cache;
detect, by the command processor, the indication of the network message in the cache; and
process, by the command processor, the network message without involving the first processor.
2 . The system as recited in claim 1 , wherein the command processor is further configured to convey the network message to a network interface unit.
3 . The system as recited in claim 2 , wherein the command processor is further configured to convey the network message directly to the network interface unit by bypassing the first processor.
4 . The system as recited in claim 1 , wherein the command processor is configured to process the network message prior to the kernel completing execution.
5 . The system as recited in claim 1 , wherein a thread of the kernel is configured to dynamically determine a target address of the network message.
6 . The system as recited in claim 5 , wherein the thread of the kernel is configured to store the network message in the cache.
7 . The system as recited in claim 1 , wherein the first processor is a central processing unit (CPU), and wherein the second processor is a graphics processing unit (GPU).
8 . A method comprising:
generating a network message within a kernel executing on a compute unit of a parallel processor; store an indication of the network message in a cache of the parallel processor; detect, by a command processor of the parallel processor, the indication of the network message in the cache; and process, by the command processor, the network message without involving a general purpose processor.
9 . The method as recited in claim 8 , further comprising conveying, by the command processor, the network message to a network interface unit.
10 . The method as recited in claim 9 , further comprising conveying, by the command processor, the network message directly to the network interface unit by bypassing the general purpose processor.
11 . The method as recited in claim 8 , further comprising processing, by the command processor, the network message prior to the kernel completing execution.
12 . The method as recited in claim 8 , further comprising dynamically determining, by a thread of the kernel, a target address of the network message.
13 . The method as recited in claim 12 , further comprising storing, by the thread of the kernel, the network message in the cache.
14 . The method as recited in claim 8 , wherein the general purpose processor is a central processing unit (CPU), and wherein the parallel processor is a graphics processing unit (GPU).
15 . An apparatus comprising:
a first processor; a second processor, wherein the second processor comprises a command processor, a plurality of compute units, and a cache; and a network interface unit; wherein the second processor is configured to:
generate a network message within a kernel executing on a compute unit;
store an indication of the network message in the cache;
detect, by the command processor, the indication of the network message in the cache; and
process, by the command processor, the network message without involving the first processor.
16 . The apparatus as recited in claim 15 , wherein the command processor is further configured to convey the network message to the network interface unit.
17 . The apparatus as recited in claim 16 , wherein the command processor is further configured to convey the network message directly to the network interface unit by bypassing the first processor.
18 . The apparatus as recited in claim 15 , wherein the command processor is configured to process the network message prior to the kernel completing execution.
19 . The apparatus as recited in claim 15 , wherein a thread of the kernel is configured to dynamically determine a target address of the network message.
20 . The apparatus as recited in claim 19 , wherein the thread of the kernel is configured to store the network message in the cache.Join the waitlist — get patent alerts
Track US2023120934A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.