Technologies for scaling deep learning training
Abstract
Technologies for artificial neural network training include a computing node with a host fabric interface that sends a message that includes one or more artificial neural network training algorithm values to another computing node in response to receipt of a request to send the message. Prior to sending the message, the host fabric interface may receive a request to quantize the message and quantize the message based on a quantization level included in the request to generate a quantized message. The quantization message includes one or more quantized values such that each quantized value has a lower precision than a corresponding artificial neural network training algorithm value. The host fabric interface then transmits the quantized message, which includes metadata indicative of the quantization level, to another computing node in response to quantization of the message for artificial neural network training. Other embodiments are described and claimed.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A compute node to send training algorithm data, the compute node comprising:
a hardware processor in circuit with a host fabric interface, the host fabric interface to include a hardware network interface controller that includes:
a quantization controller to (i) receive a first request to quantize a message, the first request indicative of a quantization level, (ii) receive a second request to compress the message, and (iii) receive a third request to send the message, the message to include one or more artificial neural network training algorithm values;
a quantizer to (i) determine the quantization level for the message in response to receipt of the third request to send the message and receipt of the first request to quantize the message, and (ii) quantize the message based on the quantization level to generate a quantized message; and
a compressor to compress the quantized message to generate a compressed quantized message in response to receipt of the second request to compress the message; and
the quantizer to transmit the compressed quantized message to a receiver computing node, the quantized message to include metadata indicative of the quantization level.
3 . The compute node of claim 2 , wherein the quantized message is to include a header, the metadata to include a field of the header of the quantized message.
4 . The compute node of claim 2 , wherein ones of the artificial neural network training algorithm values include a 32-bit floating point value.
5 . The compute node of claim 4 , wherein the quantizer is to generate the quantized message by generating one or more quantized values, the quantization level is a 16-bit level and ones of the quantized values are a fixed-point 16-bit value.
6 . The compute node of claim 4 , wherein the quantizer is to generate the quantized message by generating one or more quantized values, the quantization level is an 8-bit level and ones of the quantized values are a fixed-point 8-bit value.
7 . The computing node of claim 4 , wherein the quantizer is to generate the quantized message by generating one or more quantized values, the quantization level is a 4-bit level and ones of the quantized values are a fixed-point 4-bit value.
8 . The compute node of claim 2 , wherein the compressor is to (i) generate a bitmap including a plurality of bits, first ones of the bits indicative of whether corresponding indices of the quantized message include non-zero values, and (ii) remove second ones of the bits corresponding to zero values from the quantized message.
9 . An apparatus comprising:
at least one memory; instructions in the apparatus; and processor circuitry to execute the instructions to at least:
receive a first request to quantize a message, the first request indicative of a quantization level;
receive a second request to compress the message;
receive a third request to send the message, the message to include one or more artificial neural network training algorithm values;
determine the quantization level for the message in response to receipt of the third request to send the message and receipt of the first request to quantize the message;
quantize the message based on the quantization level to generate a quantized message;
compress the quantized message to generate a compressed quantized message in response to receipt of the second request to compress the message; and
transmit the compressed quantized message to a receiver computing node, the quantized message to include metadata indicative of the quantization level.
10 . The apparatus of claim 9 , wherein the quantized message is to include a header, the metadata to include a field of the header of the quantized message.
11 . The apparatus of claim 9 , wherein ones of the artificial neural network training algorithm values include a 32-bit floating point value.
12 . The apparatus of claim 11 , wherein the processor circuitry is to execute the instructions to generate the quantized message by generating one or more quantized values, the quantization level is a 16-bit level and ones of the quantized values are a fixed-point 16-bit value.
13 . The apparatus of claim 11 , wherein the processor circuitry is to execute the instructions to generate the quantized message by generating one or more quantized values, the quantization level is an 8-bit level and ones of the quantized values are a fixed-point 8-bit value.
14 . The apparatus of claim 11 , wherein the processor circuitry is to execute the instructions to generate the quantized message by generating one or more quantized values, the quantization level is a 4-bit level and ones of the quantized values are a fixed-point 4-bit value.
15 . The apparatus of claim 9 , wherein the processor circuitry is to execute the instructions to:
generate a bitmap including a plurality of bits, first ones of the bits indicative of whether corresponding indices of the quantized message include non-zero values; and remove second ones of the bits corresponding to zero values from the quantized message.
16 . One or more computer readable storage devices or storage disks comprising instructions that, when executed, cause at least one processor to at least:
receive a first request to quantize a message, the first request indicative of a quantization level; receive a second request to compress the message; receive a third request to send the message, the message to include one or more artificial neural network training algorithm values; determine the quantization level for the message in response to receipt of the third request to send the message and receipt of the first request to quantize the message; quantize the message based on the quantization level to generate a quantized message; compress the quantized message to generate a compressed quantized message in response to receipt of the second request to compress the message; and transmit the compressed quantized message to a receiver computing node, the quantized message to include metadata indicative of the quantization level.
17 . The one or more computer readable storage devices or storage disks of claim 16 , wherein ones of the artificial neural network training algorithm values include a 32-bit floating point value.
18 . The one or more computer readable storage devices or storage disks of claim 17 , wherein the instructions, when executed, cause the at least one processor to generate the quantized message by generating one or more quantized values, the quantization level is a 16-bit level and ones of the quantized values are a fixed-point 16-bit value.
19 . The one or more computer readable storage devices or storage disks of claim 17 , wherein the instructions, when executed, cause the at least one processor to generate the quantized message by generating one or more quantized values, the quantization level is an 8-bit level and ones of the quantized values are a fixed-point 8-bit value.
20 . The one or more computer readable storage devices or storage disks of claim 17 , wherein the instructions, when executed, cause the at least one processor to generate the quantized message by generating one or more quantized values, the quantization level is a 4-bit level and ones of the quantized values are a fixed-point 4-bit value.
21 . The one or more computer readable storage devices or storage disks of claim 16 , wherein the instructions, when executed, cause the at least one processor to:
generate a bitmap including a plurality of bits, first ones of the bits indicative of whether corresponding indices of the quantized message include non-zero values; and remove second ones of the bits corresponding to zero values from the quantized message.Join the waitlist — get patent alerts
Track US2021342692A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.