US2015302296A1PendingUtilityA1

Plastic action-selection networks for neuromorphic hardware

Assignee: HRL LAB LLCPriority: Dec 3, 2012Filed: May 16, 2013Published: Oct 22, 2015
Est. expiryDec 3, 2032(~6.3 yrs left)· nominal 20-yr term from priority
G06N 3/092G06N 3/049G06N 20/00G06N 3/0499G06N 3/08G06N 3/04
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A neural model for reinforcement-learning and for action-selection includes a plurality of channels, a population of input neurons in each of the channels, a population of output neurons in each of the channels, each population of input neurons in each of the channels coupled to each population of output neurons in each of the channels, and a population of reward neurons in each of the channels. Each channel of a population of reward neurons receives input from an environmental input, and is coupled only to output neurons in a channel that the reward neuron is part of. If the environmental input for a channel is positive, the corresponding channel of a population of output neurons are rewarded and have their responses reinforced, otherwise the corresponding channel of a population of output neurons are punished and have their responses attenuated.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A neural model for reinforcement-learning and for action-selection comprising:
 a plurality of channels;   a population of input neurons in each of the channels;   a population of output neurons in each of the channels, each population of input neurons in each of the channels coupled to each population of output neurons in each of the channels; and   a population of reward neurons in each of the channels, wherein each population of reward neurons receives input from an environmental input, and wherein each channel of reward neurons is coupled only to output neurons in a channel that the reward neuron is part of;   wherein if the environmental input for a channel is positive, the corresponding channel of a population of output neurons are rewarded and have their responses reinforced; and   wherein if the environmental input for a channel is negative, the corresponding channel of a population of output neurons are punished and have their responses attenuated.   
     
     
         2 . The neural model of  claim 1  wherein each population of output neurons in each of the channels are coupled to each population of input neurons in each of the channels by a synapse having spike-timing dependent plasticity behaving according to
     g   eff   →g   eff   +g   effmax   F (Δ t )
 
 where 
 
       
         
           
             
               
                 Δ 
                  
                 
                     
                 
                  
                 t 
               
               = 
               
                 
                   t 
                   pre 
                 
                 - 
                 
                   t 
                   post 
                 
               
             
           
         
         
           
             
               
                 F 
                  
                 
                   ( 
                   
                     Δ 
                      
                     
                         
                     
                      
                     t 
                   
                   ) 
                 
               
               = 
               
                 { 
                 
                   
                     
                       
                         
                           
                             
                               A 
                               + 
                             
                              
                             
                                
                               
                                 ( 
                                 
                                   
                                     Δ 
                                      
                                     
                                         
                                     
                                      
                                     t 
                                   
                                   
                                     τ 
                                     + 
                                   
                                 
                                 ) 
                               
                             
                           
                         
                       
                       
                         
                           
                             
                               A 
                               - 
                             
                              
                             
                                
                               
                                 ( 
                                 
                                   
                                     Δ 
                                      
                                     
                                         
                                     
                                      
                                     t 
                                   
                                   
                                     τ 
                                     - 
                                   
                                 
                                 ) 
                               
                             
                           
                         
                       
                     
                      
                     
                       
 
                     
                      
                     if 
                      
                     
                         
                     
                      
                     
                       ( 
                       
                         
                           g 
                           eff 
                         
                         < 
                         0 
                       
                       ) 
                     
                      
                     
                         
                     
                      
                     then 
                      
                     
                         
                     
                      
                     
                       g 
                       eff 
                     
                   
                   -> 
                   
                     
                       0 
                        
                       
                         
 
                       
                        
                       if 
                        
                       
                           
                       
                        
                       
                         ( 
                         
                           g 
                           > 
                           
                             g 
                             effmax 
                           
                         
                         ) 
                       
                        
                       
                           
                       
                        
                       then 
                        
                       
                         
                             
                         
                          
                         
                             
                         
                       
                        
                       
                         g 
                         eff 
                       
                     
                     -> 
                     
                       
                         g 
                         effmax 
                       
                       . 
                     
                   
                 
               
             
           
         
       
     
     
         3 . The neural model of  claim 1  wherein each population of input neurons, each population of output neurons, and each population of reward neurons are modeled with a Leaky-Integrate and Fire (LIF) model behaving according to 
       
         
           
             
               
                 
                   
                     
                       
                         C 
                         m 
                       
                        
                       
                         
                            
                           V 
                         
                         
                            
                           t 
                         
                       
                     
                     = 
                     
                       
                         - 
                         
                           
                             g 
                             leak 
                           
                            
                           
                             ( 
                             
                               V 
                               - 
                               
                                 E 
                                 rest 
                               
                             
                             ) 
                           
                         
                       
                       + 
                       
                         I 
                         . 
                       
                     
                   
                 
                 
                   
                       
                   
                 
               
             
           
         
         where
 Cm is the membrane capacitance, 
 I is the sum of external and synaptic currents, 
 
         gleak conductance of the leak channels, and 
         Erest is the reversal potential for that particular class of synapse. 
       
     
     
         4 . The neural model of  claim 1  wherein the populations of input neurons are connected with equal probability and equal conductance to all of the populations of output neurons. 
     
     
         5 . The neural model of  claim 1  wherein the populations of input neurons are connected randomly to the populations of output neurons. 
     
     
         6 . The neural model of  claim 1  wherein the neural model is implemented with a memristor based neuromorphic processor. 
     
     
         7 . A neural model for reinforcement-learning and for action-selection comprising:
 a plurality of channels;   a population of input neurons in each of the channels;   a population of output neurons in each of the channels, each population of input neurons in each of the channels coupled to each population of output neurons in each of the channels;   a population of reward neurons in each of the channels, wherein each population of reward neurons receives input from an environmental input, and wherein each channel of reward neurons is coupled only to output neurons in a channel that the reward neuron is part of; and   a population of inhibition neurons in each of the channels, wherein each population of inhibition neurons receive an input from a population of output neurons in a same channel that the population of inhibition neurons is part of, and wherein a population of inhibition neurons in a channel has an output to output neurons in every other channel except the channel of which the inhibition neurons are part of;   wherein if the environmental input to a population of reward neurons for a channel is positive, the corresponding channel of a population of output neurons are rewarded and have their responses reinforced; and   wherein if the environmental input to a population of reward neurons for a channel is negative, the corresponding channel of a population of output neurons are punished and have their responses attenuated.   
     
     
         8 . The neural model of  claim 7  wherein:
 each population of output neurons in each of the channels are coupled to each population of input neurons in each of the channels by a synapse having spike-timing dependent plasticity; 
 each channel of reward neurons is coupled to output neurons by a synapse having spike-timing dependent plasticity; 
 the input to each population of inhibition neurons from a population of output neurons in a same channel that the population of inhibition neurons is part of is by a synapse having spike-timing dependent plasticity; and 
 the output from each population of inhibition neurons in a channel is coupled to output neurons in every other channel except the channel of which the inhibition neurons are part of by a synapse having spike-timing dependent plasticity; 
 wherein the spike-timing dependent plasticity of each synapse behaves according to
     g   eff   →g   eff   +g   effmax   F (Δ t )
 
 
 where 
 
       
         
           
             
               
                 Δ 
                  
                 
                     
                 
                  
                 t 
               
               = 
               
                 
                   t 
                   pre 
                 
                 - 
                 
                   t 
                   post 
                 
               
             
           
         
         
           
             
               
                 F 
                  
                 
                   ( 
                   
                     Δ 
                      
                     
                         
                     
                      
                     t 
                   
                   ) 
                 
               
               = 
               
                 { 
                 
                   
                     
                       
                         
                           
                             
                               A 
                               + 
                             
                              
                             
                                
                               
                                 ( 
                                 
                                   
                                     Δ 
                                      
                                     
                                         
                                     
                                      
                                     t 
                                   
                                   
                                     τ 
                                     + 
                                   
                                 
                                 ) 
                               
                             
                           
                         
                       
                       
                         
                           
                             
                               A 
                               - 
                             
                              
                             
                                
                               
                                 ( 
                                 
                                   
                                     Δ 
                                      
                                     
                                         
                                     
                                      
                                     t 
                                   
                                   
                                     τ 
                                     - 
                                   
                                 
                                 ) 
                               
                             
                           
                         
                       
                     
                      
                     
                       
 
                     
                      
                     if 
                      
                     
                         
                     
                      
                     
                       ( 
                       
                         
                           g 
                           eff 
                         
                         < 
                         0 
                       
                       ) 
                     
                      
                     
                         
                     
                      
                     then 
                      
                     
                         
                     
                      
                     
                       g 
                       eff 
                     
                   
                   -> 
                   
                     
                       0 
                        
                       
                         
 
                       
                        
                       if 
                        
                       
                           
                       
                        
                       
                         ( 
                         
                           g 
                           > 
                           
                             g 
                             effmax 
                           
                         
                         ) 
                       
                        
                       
                           
                       
                        
                       then 
                        
                       
                         
                             
                         
                          
                         
                             
                         
                       
                        
                       
                         g 
                         eff 
                       
                     
                     -> 
                     
                       
                         g 
                         effmax 
                       
                       . 
                     
                   
                 
               
             
           
         
       
     
     
         9 . The neural model of  claim 7  wherein each population of input neurons, each population of output neurons, each population of reward neurons, and each population of inhibition neurons are modeled with a Leaky-Integrate and Fire (LIF) model behaving according to 
       
         
           
             
               
                 
                   
                     
                       
                         C 
                         m 
                       
                        
                       
                         
                            
                           V 
                         
                         
                            
                           t 
                         
                       
                     
                     = 
                     
                       
                         - 
                         
                           
                             g 
                             leak 
                           
                            
                           
                             ( 
                             
                               V 
                               - 
                               
                                 E 
                                 rest 
                               
                             
                             ) 
                           
                         
                       
                       + 
                       
                         I 
                         . 
                       
                     
                   
                 
                 
                   
                       
                   
                 
               
             
           
         
         where
 Cm is the membrane capacitance, 
 I is the sum of external and synaptic currents, 
 
         gleak conductance of the leak channels, and 
         Erest is the reversal potential for that particular class of synapse. 
       
     
     
         10 . The neural model of  claim 7  wherein the populations of input neurons are connected with equal probability and equal conductance to all of the populations of output neurons. 
     
     
         11 . The neural model of  claim 7  wherein the populations of input neurons are connected randomly to the populations of output neurons. 
     
     
         12 . The neural model of  claim 7  wherein as a response increases from output neurons of a channel of which a population of inhibition neurons is part of, the inhibition neurons inhibit the responses from populations of output neurons in every other channel. 
     
     
         13 . The neural model of  claim 7  wherein the neural model is implemented with a memristor based neuromorphic processor. 
     
     
         14 . A basal ganglia neural network model comprising:
 a plurality of channels;   a population of cortex neurons in each of the channels;   a population of striatum neurons in each of the channels, each population of striatum neurons in each of the channels coupled to each population of cortex neurons in each of the channels;   a population of reward neurons in each of the channels, wherein each population of reward neurons receives input from an environmental input, and wherein each channel of reward neurons is coupled only to striatum neurons in a channel that the reward neuron is part of; and   a population of Substantia Nigra pars reticulata (SNr) neurons in each of the channels, wherein each population of SNr neurons is coupled only to a population of striatum neurons in a channel that the SNr neurons are part of;   wherein if the environmental input to a population of reward neurons for a channel is positive, the corresponding channel of a population of striatum neurons are rewarded and have their responses reinforced;   wherein if the environmental input to a population of reward neurons for a channel is negative, the corresponding channel of a population of striatum neurons are punished and have their responses attenuated; and   wherein each population of SNr neurons is tonically active and is suppressed by inhibitory afferents of striatum neurons in a channel that the SNr neurons are part of.   
     
     
         15 . The basal ganglia neural network model of  claim 14  wherein:
 each population of cortex neurons in each of the channels are coupled to each population of striatum neurons in each of the channels by a synapse having spike-timing dependent plasticity; 
 each population of striatum neurons in a channel are coupled to striatum neurons in every other channel by a synapse having spike-timing dependent plasticity; 
 each channel of reward neurons is coupled to a population of striatum neurons in a same channel by a synapse having spike-timing dependent plasticity; 
 each population of SNr neurons is coupled to a population of striatum neurons in a same channel that the population of SNr neurons is part of by a synapse having spike-timing dependent plasticity; and 
 wherein the spike-timing dependent plasticity of each synapse behaves according to
     g   eff   →g   eff   +g   effmax   F (Δ t )
 
 
 where 
 
       
         
           
             
               
                 Δ 
                  
                 
                     
                 
                  
                 t 
               
               = 
               
                 
                   t 
                   pre 
                 
                 - 
                 
                   t 
                   post 
                 
               
             
           
         
         
           
             
               
                 F 
                  
                 
                   ( 
                   
                     Δ 
                      
                     
                         
                     
                      
                     t 
                   
                   ) 
                 
               
               = 
               
                 { 
                 
                   
                     
                       
                         
                           
                             
                               A 
                               + 
                             
                              
                             
                                
                               
                                 ( 
                                 
                                   
                                     Δ 
                                      
                                     
                                         
                                     
                                      
                                     t 
                                   
                                   
                                     τ 
                                     + 
                                   
                                 
                                 ) 
                               
                             
                           
                         
                       
                       
                         
                           
                             
                               A 
                               - 
                             
                              
                             
                                
                               
                                 ( 
                                 
                                   
                                     Δ 
                                      
                                     
                                         
                                     
                                      
                                     t 
                                   
                                   
                                     τ 
                                     - 
                                   
                                 
                                 ) 
                               
                             
                           
                         
                       
                     
                      
                     
                       
 
                     
                      
                     if 
                      
                     
                         
                     
                      
                     
                       ( 
                       
                         
                           g 
                           eff 
                         
                         < 
                         0 
                       
                       ) 
                     
                      
                     
                         
                     
                      
                     then 
                      
                     
                         
                     
                      
                     
                       g 
                       eff 
                     
                   
                   -> 
                   
                     
                       0 
                        
                       
                         
 
                       
                        
                       if 
                        
                       
                           
                       
                        
                       
                         ( 
                         
                           g 
                           > 
                           
                             g 
                             effmax 
                           
                         
                         ) 
                       
                        
                       
                           
                       
                        
                       then 
                        
                       
                         
                             
                         
                          
                         
                             
                         
                       
                        
                       
                         g 
                         eff 
                       
                     
                     -> 
                     
                       
                         g 
                         effmax 
                       
                       . 
                     
                   
                 
               
             
           
         
       
     
     
         16 . The basal ganglia neural network model of  claim 14  wherein each population of cortex neurons, each population of striatum neurons, each population of reward neurons, and each population of SNr neurons are modeled with a Leaky-Integrate and Fire (LIF) model behaving according to 
       
         
           
             
               
                 
                   
                     
                       
                         C 
                         m 
                       
                        
                       
                         
                            
                           V 
                         
                         
                            
                           t 
                         
                       
                     
                     = 
                     
                       
                         - 
                         
                           
                             g 
                             leak 
                           
                            
                           
                             ( 
                             
                               V 
                               - 
                               
                                 E 
                                 rest 
                               
                             
                             ) 
                           
                         
                       
                       + 
                       
                         I 
                         . 
                       
                     
                   
                 
                 
                   
                       
                   
                 
               
             
           
         
         where
 Cm is the membrane capacitance, 
 I is the sum of external and synaptic currents, 
 
         gleak conductance of the leak channels, and 
         Erest is the reversal potential for that particular class of synapse. 
       
     
     
         17 . The basal ganglia neural network model of  claim 14  wherein the populations of cortex neurons are connected with equal probability and equal conductance to all of the populations of striatum neurons. 
     
     
         18 . The basal ganglia neural network model of  claim 14  wherein the populations of cortex neurons are connected randomly to the populations of striatum neurons. 
     
     
         19 . The basal ganglia neural network model of  claim 14  wherein a Poisson random excitation is injected into the populations of SNr neurons. 
     
     
         20 . The basal ganglia neural network model of  claim 14  wherein uniform random noise is injected into the populations of SNr neurons. 
     
     
         21 . The basal ganglia neural network model of  claim 14  wherein the basal ganglia neural network model is implemented with a memristor based neuromorphic processor.

Join the waitlist — get patent alerts

Track US2015302296A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.