US2021295136A1PendingUtilityA1
Improvement of Prediction Performance Using Asymmetric Tanh Activation Function
Est. expiryOct 29, 2038(~12.3 yrs left)· nominal 20-yr term from priority
Inventors:Yong Hee Han
G06N 3/048G06N 3/045G06N 3/088G06N 3/0985G06N 3/09G06N 3/0455G06N 3/0499G06N 3/08G06N 3/0481
47
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The present disclosure in at least one aspect provides an asymmetric hyperbolic tangent (tanh) function which can be used as an activation function irrespective of the structure of a neural network. The activation function provided limits an output range thereof to between a maximum value and a minimum value of a variable to be predicted. The activation function provided is suitable for a regression problem which requires the prediction of a wide range of real values depending on input data.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, implemented by a computer, for processing data representing an actual phenomenon by using a neural network configured to model an actual data pattern, the method comprising:
at each node of an output layer of the neural network, computing weighted sum of input values, the input values at each node of the output layer of the neural network being output values from nodes of a last hidden layer of at least one hidden layer of the neural network; and at each nodes of the output layer of the neural network, applying a nonlinear activation function to the weighted sum of the input values to generate output value, wherein the nonlinear activation function has an output range with an upper limit and a lower limit that are respectively bounded by a maximum value and a minimum value of a variable to be predicted at a relevant node of the output layer of the neural network.
2 . The method of claim 1 , wherein the nonlinear activation function is expressed by an equation:
f
(
x
)
=
{
tanh
(
x
max
/
s
)
×
max
if
x
>
0
tanh
(
x
min
/
s
)
×
min
else
,
wherein
x is a weighted sum of the input values at the relevant node of the output layer of the neural network, max and min are respectively the maximum value and the minimum value of the variable to be predicted at the relevant node of the output layer of the neural network, and s is a parameter that adjusts a derivative of the nonlinear activation function.
3 . The method of claim 2 , wherein the variable to be predicted at the relevant node of the output layer of the neural network is data inputted to a relevant node of an input layer of the neural network.
4 . The method of claim 2 , wherein the parameter is set to a hyper-parameter or to be learned from training data.
5 . The method of claim 1 , wherein the nonlinear activation function is expressed by an equation:
f
(
x
)
=
{
tanh
(
x
max
)
×
max
if
x
>
0
tanh
(
x
min
)
×
min
else
,
wherein
x is a weighted sum of the input values at the relevant node of the output layer, and max and min are respectively the maximum value and the minimum value of the variable to be predicted at the relevant node of the output layer of the neural network.
6 . The method of claim 1 , further comprising:
detecting anomaly data out of the data representing the actual phenomenon based on a difference between data inputted to each node of an input layer of the neural network and an output value generated at each node of the output layer of the neural network.
7 . The method of claim 1 , further comprising:
utilizing output values from nodes of any hidden layer of the at least one hidden layer of the neural network as compressed representations of data inputted to nodes of an input layer of the neural network.
8 . An apparatus for processing data representing an actual phenomenon by using a neural network configured to model an actual data pattern, the apparatus comprising:
at least one processor; and at least one memory in which instructions are recorded, wherein the instructions cause, when executed in the processor, the processor to perform steps comprising:
at each node of an output layer of the neural network, computing weighted sum of input values, the input values at each node of the output layer of the neural network being output values from nodes of a last hidden layer of at least one hidden layer of the neural network; and
at each node of the output layer of the neural network, applying a nonlinear activation function to the weighted sum of the input values to generate output value, wherein
the nonlinear activation function has an output range with an upper limit and a lower limit that are respectively bounded by a maximum value and a minimum value of a variable to be predicted at a relevant node of the output layer of the neural network.
9 . The apparatus of claim 8 , wherein the nonlinear activation function is expressed by an equation:
f
(
x
)
=
{
tanh
(
x
max
/
s
)
×
max
if
x
>
0
tanh
(
x
min
/
s
)
×
min
else
,
wherein
x is a weighted sum of the input values at the relevant node of the output layer, max and min are respectively the maximum value and the minimum value of the variable to be predicted at the relevant node of the output layer of the neural network, and s is a parameter that adjusts a derivative of the nonlinear activation function.
10 . The apparatus of claim 8 , wherein the nonlinear activation function is expressed by an equation:
f
(
x
)
=
{
tanh
(
x
max
)
×
max
if
x
>
0
tanh
(
x
min
)
×
min
else
,
wherein
x is a weighted sum of the input values at the relevant node of the output layer, and max and min are respectively the maximum value and the minimum value of the variable to be predicted at the relevant node of the output layer of the neural network.
11 . An apparatus for performing a neural network operation for a neural network configured to model an actual data pattern to process data representing an actual phenomenon, the apparatus comprising:
a weighted sum operation unit configured to receive input values and weights for nodes of an output layer of the neural network and to generate a plurality of weighted sums for the nodes of the output layer of the neural network based on the input values and the weights that are received, the input values at each node of the output layer of the neural network being output values from nodes of a last hidden layer of at least one hidden layer of the neural network; and an output operation unit configured to apply an activation function to weighted sums of the respective nodes of the output layer of the neural network to generate output values for the respective nodes of the output layer of the neural network, wherein the nonlinear activation function has an output range with an upper limit and a lower limit that are respectively bounded by a maximum value and a minimum value of a variable to be predicted at a relevant node of the output layer of the neural network.
12 . The apparatus of claim 11 , wherein the nonlinear activation function is expressed by an equation:
f
(
x
)
=
{
tanh
(
x
max
/
s
)
×
max
if
x
>
0
tanh
(
x
min
/
s
)
×
min
else
,
wherein
x is a weighted sum of the input values at the relevant node of the output layer of the neural network, max and min are respectively the maximum value and the minimum value of the variable to be predicted at the relevant node of the output layer of the neural network, and s is a parameter that adjusts a derivative of the nonlinear activation function.
13 . The apparatus of claim 11 , wherein the nonlinear activation function is expressed by an equation:
f
(
x
)
=
{
tanh
(
x
max
)
×
max
if
x
>
0
tanh
(
x
min
)
×
min
else
,
wherein
x is a weighted sum of the input values at the relevant node of the output layer, and max and min are respectively the maximum value and the minimum value of the variable to be predicted at the relevant node of the output layer of the neural network.Join the waitlist — get patent alerts
Track US2021295136A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.