Plastic action-selection networks for neuromorphic hardware
Abstract
A neural model for reinforcement-learning and for action-selection includes a plurality of channels, a population of input neurons in each of the channels, a population of output neurons in each of the channels, each population of input neurons in each of the channels coupled to each population of output neurons in each of the channels, and a population of reward neurons in each of the channels. Each channel of a population of reward neurons receives input from an environmental input, and is coupled only to output neurons in a channel that the reward neuron is part of. If the environmental input for a channel is positive, the corresponding channel of a population of output neurons are rewarded and have their responses reinforced, otherwise the corresponding channel of a population of output neurons are punished and have their responses attenuated.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A neural model for reinforcement-learning and for action-selection comprising:
a plurality of channels; a population of input neurons in each of the channels; a population of output neurons in each of the channels, each population of input neurons in each of the channels coupled to each population of output neurons in each of the channels; and a population of reward neurons in each of the channels, wherein each population of reward neurons receives input from an environmental input, and wherein each channel of reward neurons is coupled only to output neurons in a channel that the reward neuron is part of; wherein if the environmental input for a channel is positive, the corresponding channel of a population of output neurons are rewarded and have their responses reinforced; and wherein if the environmental input for a channel is negative, the corresponding channel of a population of output neurons are punished and have their responses attenuated.
2 . The neural model of claim 1 wherein each population of output neurons in each of the channels are coupled to each population of input neurons in each of the channels by a synapse having spike-timing dependent plasticity behaving according to
g eff →g eff +g effmax F (Δ t )
where
Δ
t
=
t
pre
-
t
post
F
(
Δ
t
)
=
{
A
+
(
Δ
t
τ
+
)
A
-
(
Δ
t
τ
-
)
if
(
g
eff
<
0
)
then
g
eff
->
0
if
(
g
>
g
effmax
)
then
g
eff
->
g
effmax
.
3 . The neural model of claim 1 wherein each population of input neurons, each population of output neurons, and each population of reward neurons are modeled with a Leaky-Integrate and Fire (LIF) model behaving according to
C
m
V
t
=
-
g
leak
(
V
-
E
rest
)
+
I
.
where
Cm is the membrane capacitance,
I is the sum of external and synaptic currents,
gleak conductance of the leak channels, and
Erest is the reversal potential for that particular class of synapse.
4 . The neural model of claim 1 wherein the populations of input neurons are connected with equal probability and equal conductance to all of the populations of output neurons.
5 . The neural model of claim 1 wherein the populations of input neurons are connected randomly to the populations of output neurons.
6 . The neural model of claim 1 wherein the neural model is implemented with a memristor based neuromorphic processor.
7 . A neural model for reinforcement-learning and for action-selection comprising:
a plurality of channels; a population of input neurons in each of the channels; a population of output neurons in each of the channels, each population of input neurons in each of the channels coupled to each population of output neurons in each of the channels; a population of reward neurons in each of the channels, wherein each population of reward neurons receives input from an environmental input, and wherein each channel of reward neurons is coupled only to output neurons in a channel that the reward neuron is part of; and a population of inhibition neurons in each of the channels, wherein each population of inhibition neurons receive an input from a population of output neurons in a same channel that the population of inhibition neurons is part of, and wherein a population of inhibition neurons in a channel has an output to output neurons in every other channel except the channel of which the inhibition neurons are part of; wherein if the environmental input to a population of reward neurons for a channel is positive, the corresponding channel of a population of output neurons are rewarded and have their responses reinforced; and wherein if the environmental input to a population of reward neurons for a channel is negative, the corresponding channel of a population of output neurons are punished and have their responses attenuated.
8 . The neural model of claim 7 wherein:
each population of output neurons in each of the channels are coupled to each population of input neurons in each of the channels by a synapse having spike-timing dependent plasticity;
each channel of reward neurons is coupled to output neurons by a synapse having spike-timing dependent plasticity;
the input to each population of inhibition neurons from a population of output neurons in a same channel that the population of inhibition neurons is part of is by a synapse having spike-timing dependent plasticity; and
the output from each population of inhibition neurons in a channel is coupled to output neurons in every other channel except the channel of which the inhibition neurons are part of by a synapse having spike-timing dependent plasticity;
wherein the spike-timing dependent plasticity of each synapse behaves according to
g eff →g eff +g effmax F (Δ t )
where
Δ
t
=
t
pre
-
t
post
F
(
Δ
t
)
=
{
A
+
(
Δ
t
τ
+
)
A
-
(
Δ
t
τ
-
)
if
(
g
eff
<
0
)
then
g
eff
->
0
if
(
g
>
g
effmax
)
then
g
eff
->
g
effmax
.
9 . The neural model of claim 7 wherein each population of input neurons, each population of output neurons, each population of reward neurons, and each population of inhibition neurons are modeled with a Leaky-Integrate and Fire (LIF) model behaving according to
C
m
V
t
=
-
g
leak
(
V
-
E
rest
)
+
I
.
where
Cm is the membrane capacitance,
I is the sum of external and synaptic currents,
gleak conductance of the leak channels, and
Erest is the reversal potential for that particular class of synapse.
10 . The neural model of claim 7 wherein the populations of input neurons are connected with equal probability and equal conductance to all of the populations of output neurons.
11 . The neural model of claim 7 wherein the populations of input neurons are connected randomly to the populations of output neurons.
12 . The neural model of claim 7 wherein as a response increases from output neurons of a channel of which a population of inhibition neurons is part of, the inhibition neurons inhibit the responses from populations of output neurons in every other channel.
13 . The neural model of claim 7 wherein the neural model is implemented with a memristor based neuromorphic processor.
14 . A basal ganglia neural network model comprising:
a plurality of channels; a population of cortex neurons in each of the channels; a population of striatum neurons in each of the channels, each population of striatum neurons in each of the channels coupled to each population of cortex neurons in each of the channels; a population of reward neurons in each of the channels, wherein each population of reward neurons receives input from an environmental input, and wherein each channel of reward neurons is coupled only to striatum neurons in a channel that the reward neuron is part of; and a population of Substantia Nigra pars reticulata (SNr) neurons in each of the channels, wherein each population of SNr neurons is coupled only to a population of striatum neurons in a channel that the SNr neurons are part of; wherein if the environmental input to a population of reward neurons for a channel is positive, the corresponding channel of a population of striatum neurons are rewarded and have their responses reinforced; wherein if the environmental input to a population of reward neurons for a channel is negative, the corresponding channel of a population of striatum neurons are punished and have their responses attenuated; and wherein each population of SNr neurons is tonically active and is suppressed by inhibitory afferents of striatum neurons in a channel that the SNr neurons are part of.
15 . The basal ganglia neural network model of claim 14 wherein:
each population of cortex neurons in each of the channels are coupled to each population of striatum neurons in each of the channels by a synapse having spike-timing dependent plasticity;
each population of striatum neurons in a channel are coupled to striatum neurons in every other channel by a synapse having spike-timing dependent plasticity;
each channel of reward neurons is coupled to a population of striatum neurons in a same channel by a synapse having spike-timing dependent plasticity;
each population of SNr neurons is coupled to a population of striatum neurons in a same channel that the population of SNr neurons is part of by a synapse having spike-timing dependent plasticity; and
wherein the spike-timing dependent plasticity of each synapse behaves according to
g eff →g eff +g effmax F (Δ t )
where
Δ
t
=
t
pre
-
t
post
F
(
Δ
t
)
=
{
A
+
(
Δ
t
τ
+
)
A
-
(
Δ
t
τ
-
)
if
(
g
eff
<
0
)
then
g
eff
->
0
if
(
g
>
g
effmax
)
then
g
eff
->
g
effmax
.
16 . The basal ganglia neural network model of claim 14 wherein each population of cortex neurons, each population of striatum neurons, each population of reward neurons, and each population of SNr neurons are modeled with a Leaky-Integrate and Fire (LIF) model behaving according to
C
m
V
t
=
-
g
leak
(
V
-
E
rest
)
+
I
.
where
Cm is the membrane capacitance,
I is the sum of external and synaptic currents,
gleak conductance of the leak channels, and
Erest is the reversal potential for that particular class of synapse.
17 . The basal ganglia neural network model of claim 14 wherein the populations of cortex neurons are connected with equal probability and equal conductance to all of the populations of striatum neurons.
18 . The basal ganglia neural network model of claim 14 wherein the populations of cortex neurons are connected randomly to the populations of striatum neurons.
19 . The basal ganglia neural network model of claim 14 wherein a Poisson random excitation is injected into the populations of SNr neurons.
20 . The basal ganglia neural network model of claim 14 wherein uniform random noise is injected into the populations of SNr neurons.
21 . The basal ganglia neural network model of claim 14 wherein the basal ganglia neural network model is implemented with a memristor based neuromorphic processor.Join the waitlist — get patent alerts
Track US2015302296A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.