Persistent kernel for graphics processing unit direct memory access network packet processing
Abstract
A graphics processing unit may, in accordance with a kernel, determine that at least a first packet is written to a memory buffer of the graphics processing unit by a network interface card via a direct memory access, process the at least the first packet in accordance with the kernel, and provide a first notification to a central processing unit that the at least the first packet is processed in accordance with the kernel. The graphics processing unit may further determine that at least a second packet is written to the memory buffer by the network interface card via the direct memory access, process the at least the second packet in accordance with the kernel, where the kernel comprises a persistent kernel, and provide a second notification to the central processing unit that the at least the second packet is processed in accordance with the kernel.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
sending, by a central processing unit, an instruction to a graphics processing unit to allocate a memory buffer for direct memory access; instantiating, by the central processing unit, a kernel at the graphics processing unit; receiving, by the central processing unit, a first confirmation having a starting address from the graphics processing unit that the memory buffer has been allocated; sending, by the central processing unit, a message to a computing device to authorize the computing device to directly write at least a first packet to the memory buffer of the graphics processing unit via a direct memory access; and sending, by the central processing unit to the graphics processing unit, a first indicator that the at least the first packet is written to the memory buffer by the computing device via the direct memory access.
2 . The method of claim 1 , wherein the confirmation further comprises a size or a length of the memory buffer.
3 . The method of claim 1 , wherein the message comprises the starting address.
4 . The method of claim 3 , wherein the message further comprises a size or a length of the memory buffer.
5 . The method of claim 1 , further comprising:
receiving, by the central processing unit, a first notification from the graphics processing unit that the at least the first packet is processed in accordance with the kernel.
6 . The method of claim 5 , further comprising:
receiving, by the central processing unit, a second notification from the graphics processing unit that at least a second packet is processed in accordance with the kernel, wherein the kernel comprises a persistent kernel.
7 . The method of claim 6 , wherein the kernel is executed by at least one streaming multiprocessor of the graphics processing unit.
8 . The method of claim 7 , wherein the kernel remains assigned to the at least one streaming multiprocessor from the processing of the at least the first packet to the processing of the at least the second packet.
9 . The method of claim 1 , further comprising:
sending, by the central processing unit, to the graphics processing unit, a second indicator that at least a second packet is written to the memory buffer by the computing device via the direct memory access.
10 . The method of claim 9 , wherein the first indicator and the second indicator comprise notification flags written to a shared memory of the graphics processing unit by the central processing unit.
11 . The method of claim 9 , wherein the first indicator and the second indicator comprise instructions from the central processing unit to a scheduler of the graphics processing unit.
12 . The method of claim 9 , wherein the at least the first packet is written to the memory buffer by the computing device via the direct memory access in a first block of packets of a defined size, and wherein the at least the second packet is written to the memory buffer by the computing device via the direct memory access in a second block of packets of the defined size.
13 . The method of claim 9 , wherein the at least the second packet is written to the memory buffer after the processing the at least the first packet via the kernel.
14 . The method of claim 1 , wherein the computing device comprises at least one of: a transceiver or a direct memory access engine.
15 . The method of claim 1 , wherein the instantiating comprises sending the kernel, by the central processing unit, to the graphics processing unit.
16 . The method of claim 15 , wherein the kernel comprises:
kernel code for execution as a plurality of threads; and instructions for arranging the plurality of threads of the kernel into at least one thread block.
17 . The method of claim 16 , wherein each of the at least one thread block is assigned to a respective streaming multiprocessor of a plurality of streaming multiprocessors of the graphics processing unit.
18 . The method of claim 1 , further comprising:
receiving, by the central processing unit from the computing device, a second confirmation that the computing device has directly written the at least the first packet to the memory buffer of the graphics processing unit via the direct memory access.
19 . A non-transitory computer-readable medium storing instructions which, when executed by a central processing unit, cause the central processing unit to perform operations, the operations comprising:
sending an instruction to a graphics processing unit to allocate a memory buffer for direct memory access; instantiating a kernel at the graphics processing unit; receiving a first confirmation having a starting address from the graphics processing unit that the memory buffer has been allocated; sending a message to a computing device to authorize the computing device to directly write at least a first packet to the memory buffer of the graphics processing unit via a direct memory access; and sending, to the graphics processing unit, a first indicator that the at least the first packet is written to the memory buffer by the computing device via the direct memory access.
20 . A device comprising:
a central processing unit; and a non-transitory computer-readable medium storing instructions which, when executed by the graphics processing unit, cause the graphics processing unit to perform operations, the operations comprising:
sending an instruction to a graphics processing unit to allocate a memory buffer for direct memory access;
instantiating a kernel at the graphics processing unit;
receiving a first confirmation having a starting address from the graphics processing unit that the memory buffer has been allocated;
sending a message to a computing device to authorize the computing device to directly write at least a first packet to the memory buffer of the graphics processing unit via a direct memory access; and
sending, to the graphics processing unit, a first indicator that the at least the first packet is written to the memory buffer by the computing device via the direct memory access.Join the waitlist — get patent alerts
Track US2022261367A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.