US2022107783A1PendingUtilityA1
Machine learning training architecture for programmable devices
Est. expiryMar 27, 2039(~12.7 yrs left)· nominal 20-yr term from priority
G06N 3/048G06F 17/16H03M 7/24G06N 3/063G06F 7/501G06N 20/00G06N 3/084G06F 7/5443G06N 3/02G06F 7/485G06F 2207/4824G06F 9/30014G06F 7/4876
67
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A programmable device may be configured to support machine learning training operations using matrix multiplication circuitry. In some embodiments, the multiplication is implemented on a systolic array. The systolic array includes an array of processing elements, each of which includes hybrid floating-point dot-product circuitry.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A programmable logic device (PLD), comprising:
machine learning training circuitry, configured to train a neural network, comprising:
a pipeline, comprising:
a first stage circuitry configured to load a first matrix and a second matrix from off-chip memory;
a second stage configured to perform matrix multiplication of the first matrix and the second matrix; and
a third stage configured to load a result of the second stage matrix multiplication to the off-chip memory.
2 . The programmable logic device of claim 1 , wherein the first stage circuitry comprises:
a first load circuit configured to load the first matrix on-chip from the off-chip memory; and a second load circuit configured to load the second matrix on-chip from the off-chip memory.
3 . The programmable logic device of claim 2 , configured to:
reduce memory traffic by performing one or more transpositions, activations functions or both within the pipeline to: mutate the first matrix, the second matrix, or both loaded from off-chip memory, wherein the first load circuit, the second load circuit, or both; mutate the result of the second stage matrix multiplication prior to loading the result to the off-chip memory; or both.
4 . The programmable logic device of claim 1 , configured to:
perform the training of the neural network, comprising a multilayer perception, via a two-pass execution by the PLD performed at successive layers of the multilayer perception, comprising:
a forward pass that performs the matrix multiplication of the first matrix and the second matrix, wherein the first matrix comprises a current weight matrix and the second matrix comprises a prior layer's output passed through an activation function; and
a backward pass that computes a gradient of the activation function and determines changes to be applied to the current weight matrix.
5 . The programmable logic device of claim 4 , comprising:
a stochastic gradient descent circuit configured to implement the backward pass via stochastic gradient descent.
6 . The programmable logic device of claim 1 , configured to enhance off-memory matrix access, by loading the second stage matrix multiplication to the off-chip memory in an ordered manner, such that one or more sequences of consecutive addresses may be grouped in bursts for joint retrieval.
7 . The programmable logic device of claim 1 , comprising one or more systolic arrays, wherein the second stage is configured to perform the matrix multiplication of the first matrix and the second matrix using the one or more systolic arrays.
8 . The programmable logic device of claim 7 , wherein the one or more systolic arrays comprise:
one or more processing elements; and control logic for coordinating the one or more processing elements.
9 . The programmable logic device of claim 8 , comprising a row feeder, wherein the one or more processing elements comprise a row of processing elements fed with at least a portion of the first matrix via the row feeder.
10 . The programmable logic device of claim 8 , comprising a column feeder, wherein the one or more processing elements comprise a column of processing elements fed with at least a portion of the second matrix via the column feeder.
11 . The programmable logic device of claim 8 , wherein the one or more processing elements comprise:
a hybrid floating-point dot-product circuitry comprising both a hard floating-point multiplier and a soft floating-point multiplier.
12 . The programmable logic device of claim 11 , comprising:
one or more delay registers between circuitry in the first stage and circuitry in the second stage to counteract latency discrepancies between the hard floating-point multiplier and the soft floating-point multiplier.
13 . The programmable logic device of claim 11 , wherein the circuitry in the second stage comprises the hard floating-point multiplier.
14 . The programmable logic device of claim 11 , wherein the one or more processing elements comprise:
an accumulation storage circuit configured to:
store intermediate results of the hybrid floating-point dot-product circuitry; and
selectively feed accumulated data back as input to the hybrid floating-point dot-product circuitry.
15 . An integrated circuit, comprising:
a plurality of processing elements, arranged in rows of processing elements and columns of processing elements, wherein each of the plurality of processing elements comprises a hybrid floating-point dot-product circuitry comprising both a hard floating-point multiplier and a soft floating-point multiplier; and one or more delay registers between circuitry configured to counteract latency discrepancies between the hard floating-point multiplier and the soft floating-point multiplier.
16 . The integrated circuit of claim 15 , comprising:
a row feeder configured to feed off-chip matrix data to a corresponding row of the rows of processing elements; and a column feeder configured to feed additional off-chip matrix data to a corresponding column of the columns of processing elements.
17 . The integrated circuit of claim 15 , wherein each of the plurality of processing elements comprises an accumulation storage circuitry configured to store intermediate results of the hybrid floating-point dot-product circuitry.
18 . A programmable logic device-implemented method, comprising:
training a neural network, by:
in a first stage of a pipeline, loading a first matrix and a second matrix from off-chip memory;
in a second stage of the pipeline, performing matrix multiplication of the first matrix and the second matrix; and
in a third stage of the pipeline, loading a result of the second stage matrix multiplication to the off-chip memory.
19 . The programmable logic device-implemented method of claim 18 , comprising:
performing the training of the neural network, comprising a multilayer perception, via a two-pass execution performed at successive layers of the multilayer perception, comprising:
a forward pass that performs the matrix multiplication of the first matrix and the second matrix, wherein the first matrix comprises a current weight matrix and the second matrix comprises a prior layer's output passed through an activation function; and
a backward pass that computes a gradient of the activation function and determines changes to be applied to the current weight matrix.
20 . The programmable logic device-implemented method of claim 19 , comprising:
implementing the backward pass via stochastic gradient descent.Join the waitlist — get patent alerts
Track US2022107783A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.