US2021342692A1PendingUtilityA1

Technologies for scaling deep learning training

Assignee: INTEL CORPPriority: Apr 1, 2017Filed: May 14, 2021Published: Nov 4, 2021
Est. expiryApr 1, 2037(~10.7 yrs left)· nominal 20-yr term from priority
G06N 3/045G06N 3/098G06N 3/0464G06N 3/0495G06N 3/063G06N 3/08G06N 3/0454
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Technologies for artificial neural network training include a computing node with a host fabric interface that sends a message that includes one or more artificial neural network training algorithm values to another computing node in response to receipt of a request to send the message. Prior to sending the message, the host fabric interface may receive a request to quantize the message and quantize the message based on a quantization level included in the request to generate a quantized message. The quantization message includes one or more quantized values such that each quantized value has a lower precision than a corresponding artificial neural network training algorithm value. The host fabric interface then transmits the quantized message, which includes metadata indicative of the quantization level, to another computing node in response to quantization of the message for artificial neural network training. Other embodiments are described and claimed.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A compute node to send training algorithm data, the compute node comprising:
 a hardware processor in circuit with a host fabric interface, the host fabric interface to include a hardware network interface controller that includes:
 a quantization controller to (i) receive a first request to quantize a message, the first request indicative of a quantization level, (ii) receive a second request to compress the message, and (iii) receive a third request to send the message, the message to include one or more artificial neural network training algorithm values; 
 a quantizer to (i) determine the quantization level for the message in response to receipt of the third request to send the message and receipt of the first request to quantize the message, and (ii) quantize the message based on the quantization level to generate a quantized message; and 
 a compressor to compress the quantized message to generate a compressed quantized message in response to receipt of the second request to compress the message; and 
 the quantizer to transmit the compressed quantized message to a receiver computing node, the quantized message to include metadata indicative of the quantization level. 
   
     
     
         3 . The compute node of  claim 2 , wherein the quantized message is to include a header, the metadata to include a field of the header of the quantized message. 
     
     
         4 . The compute node of  claim 2 , wherein ones of the artificial neural network training algorithm values include a 32-bit floating point value. 
     
     
         5 . The compute node of  claim 4 , wherein the quantizer is to generate the quantized message by generating one or more quantized values, the quantization level is a 16-bit level and ones of the quantized values are a fixed-point 16-bit value. 
     
     
         6 . The compute node of  claim 4 , wherein the quantizer is to generate the quantized message by generating one or more quantized values, the quantization level is an 8-bit level and ones of the quantized values are a fixed-point 8-bit value. 
     
     
         7 . The computing node of  claim 4 , wherein the quantizer is to generate the quantized message by generating one or more quantized values, the quantization level is a 4-bit level and ones of the quantized values are a fixed-point 4-bit value. 
     
     
         8 . The compute node of  claim 2 , wherein the compressor is to (i) generate a bitmap including a plurality of bits, first ones of the bits indicative of whether corresponding indices of the quantized message include non-zero values, and (ii) remove second ones of the bits corresponding to zero values from the quantized message. 
     
     
         9 . An apparatus comprising:
 at least one memory;   instructions in the apparatus; and   processor circuitry to execute the instructions to at least:
 receive a first request to quantize a message, the first request indicative of a quantization level; 
 receive a second request to compress the message; 
 receive a third request to send the message, the message to include one or more artificial neural network training algorithm values; 
 determine the quantization level for the message in response to receipt of the third request to send the message and receipt of the first request to quantize the message; 
 quantize the message based on the quantization level to generate a quantized message; 
 compress the quantized message to generate a compressed quantized message in response to receipt of the second request to compress the message; and 
 transmit the compressed quantized message to a receiver computing node, the quantized message to include metadata indicative of the quantization level. 
   
     
     
         10 . The apparatus of  claim 9 , wherein the quantized message is to include a header, the metadata to include a field of the header of the quantized message. 
     
     
         11 . The apparatus of  claim 9 , wherein ones of the artificial neural network training algorithm values include a 32-bit floating point value. 
     
     
         12 . The apparatus of  claim 11 , wherein the processor circuitry is to execute the instructions to generate the quantized message by generating one or more quantized values, the quantization level is a 16-bit level and ones of the quantized values are a fixed-point 16-bit value. 
     
     
         13 . The apparatus of  claim 11 , wherein the processor circuitry is to execute the instructions to generate the quantized message by generating one or more quantized values, the quantization level is an 8-bit level and ones of the quantized values are a fixed-point 8-bit value. 
     
     
         14 . The apparatus of  claim 11 , wherein the processor circuitry is to execute the instructions to generate the quantized message by generating one or more quantized values, the quantization level is a 4-bit level and ones of the quantized values are a fixed-point 4-bit value. 
     
     
         15 . The apparatus of  claim 9 , wherein the processor circuitry is to execute the instructions to:
 generate a bitmap including a plurality of bits, first ones of the bits indicative of whether corresponding indices of the quantized message include non-zero values; and   remove second ones of the bits corresponding to zero values from the quantized message.   
     
     
         16 . One or more computer readable storage devices or storage disks comprising instructions that, when executed, cause at least one processor to at least:
 receive a first request to quantize a message, the first request indicative of a quantization level;   receive a second request to compress the message;   receive a third request to send the message, the message to include one or more artificial neural network training algorithm values;   determine the quantization level for the message in response to receipt of the third request to send the message and receipt of the first request to quantize the message;   quantize the message based on the quantization level to generate a quantized message;   compress the quantized message to generate a compressed quantized message in response to receipt of the second request to compress the message; and   transmit the compressed quantized message to a receiver computing node, the quantized message to include metadata indicative of the quantization level.   
     
     
         17 . The one or more computer readable storage devices or storage disks of  claim 16 , wherein ones of the artificial neural network training algorithm values include a 32-bit floating point value. 
     
     
         18 . The one or more computer readable storage devices or storage disks of  claim 17 , wherein the instructions, when executed, cause the at least one processor to generate the quantized message by generating one or more quantized values, the quantization level is a 16-bit level and ones of the quantized values are a fixed-point 16-bit value. 
     
     
         19 . The one or more computer readable storage devices or storage disks of  claim 17 , wherein the instructions, when executed, cause the at least one processor to generate the quantized message by generating one or more quantized values, the quantization level is an 8-bit level and ones of the quantized values are a fixed-point 8-bit value. 
     
     
         20 . The one or more computer readable storage devices or storage disks of  claim 17 , wherein the instructions, when executed, cause the at least one processor to generate the quantized message by generating one or more quantized values, the quantization level is a 4-bit level and ones of the quantized values are a fixed-point 4-bit value. 
     
     
         21 . The one or more computer readable storage devices or storage disks of  claim 16 , wherein the instructions, when executed, cause the at least one processor to:
 generate a bitmap including a plurality of bits, first ones of the bits indicative of whether corresponding indices of the quantized message include non-zero values; and   remove second ones of the bits corresponding to zero values from the quantized message.

Join the waitlist — get patent alerts

Track US2021342692A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.