US2023117143A1PendingUtilityA1

Efficient learning and using of topologies of neural networks in machine learning

Assignee: INTEL CORPPriority: May 5, 2017Filed: Nov 8, 2022Published: Apr 20, 2023
Est. expiryMay 5, 2037(~10.8 yrs left)· nominal 20-yr term from priority
G06N 3/08G06N 3/0475G06N 3/0464G06N 3/082G06N 3/09G06N 3/098G06N 3/0495G06N 7/01G06T 1/20G06N 3/045G06N 3/044
67
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A mechanism is described for facilitating learning and application of neural network topologies in machine learning at autonomous machines. A method of embodiments, as described herein, includes monitoring and detecting structure learning of neural networks relating to machine learning operations at a computing device having a processor, and generating a recursive generative model based on one or more topologies of one or more of the neural networks. The method may further include converting the generative model into a discriminative model.

Claims

exact text as granted — not AI-modified
1 . An apparatus comprising:
 a graphics processor to:
 learn a structure of a generative probabilistic model; 
 determine, based on the structure, a stochastic inverse of the generative probabilistic model; 
 convert the stochastic inverse into a discriminative model; 
 convert the discriminative model into a deep neural network (DNN) that is trained using labeled data; and 
 train the DNN using labeled data, wherein training the DNN comprises at least:
 setting a bit-precision of weights are set in neurons of the DNN independently from one another; and 
 performing methodological dropout of the neurons, wherein the methodological dropout is performed in accordance with a predictivity based on historical statistical data relating to the neurons. 
 
   
     
     
         2 . The apparatus of  claim 1 , wherein the generative model is unsupervised and based on unlabeled data, and wherein the discriminative model is supervised and based on labeled data, wherein the discriminative model is learned from the generative probabilistic model. 
     
     
         3 . The apparatus of  claim 1 , wherein the graphics processor is further to inverse the generative probabilistic model into multiple inverse models, wherein a bidirectional connection is added to connect latent variables having a common parent in each of the multiple inverse models to consolidate the multiple inverse models into a single inverse model. 
     
     
         4 . The apparatus of  claim 3 , wherein the graphics processor is further to convert the inverse model into the discriminative model by removing the bidirectional connection and adding a class node serving as a child node to latent leaves. 
     
     
         5 . The apparatus of  claim 1 , wherein the graphics processor is further to:
 generate parallel and sequential execution schedules for memory sharing at sub-network precision levels of the one or more of the neural networks; and   perform on-the-fly learning and updating of network topologies of the neural networks based on at least one of currently available data and historically available data relating to the topologies of the neural networks.   
     
     
         6 . The apparatus of  claim 1 , wherein the graphics processor is further to:
 facilitate at least one of an end-to-end structure learning and a sub-network structure learning; and   facilitate feature bagging or coping with large scale data by training large training sets.   
     
     
         7 . The apparatus of  claim 1 , wherein the graphics processor is co-located with an application processor on a common semiconductor package. 
     
     
         8 . A method comprising:
 learning, by a graphics processor, a structure of a generative probabilistic model;   determining, based on the structure, a stochastic inverse of the generative probabilistic model;   converting the stochastic inverse into a discriminative model;   converting the discriminative model into a deep neural network (DNN) that is trained using labeled data; and   training the DNN using labeled data, wherein training the DNN comprises at least:
 setting a bit-precision of weights are set in neurons of the DNN independently from one another; and 
 performing methodological dropout of the neurons, wherein the methodological dropout is performed in accordance with a predictivity based on historical statistical data relating to the neurons. 
   
     
     
         9 . The method of  claim 8 , wherein the generative probabilistic model is unsupervised and based on unlabeled data, and wherein the discriminative model is supervised and based on labeled data, wherein the discriminative model is learned from the generative probabilistic model. 
     
     
         10 . The method of  claim 8 , further comprising inversing the generative probabilistic model into multiple inverse models, wherein a bidirectional connection is added to connect latent variables having a common parent in each of the multiple inverse models to consolidate the multiple inverse models into a single inverse model. 
     
     
         11 . The method of  claim 10 , further comprising converting the inverse model into the discriminative model by removing the bidirectional connection and adding a class node serving as a child node to latent leaves. 
     
     
         12 . The method of  claim 8 , further comprising:
 generating parallel and sequential execution schedules for memory sharing at sub-network precision levels of the one or more of the neural networks; and   performing on-the-fly learning and updating of network topologies of the neural networks based on at least one of currently available data and historically available data relating to the topologies of the neural networks.   
     
     
         13 . The method of  claim 8 , further comprising:
 facilitating at least one of an end-to-end structure learning and a sub-network structure learning; and   facilitating feature bagging or coping with large scale data by training large training sets.   
     
     
         14 . The method of  claim 8 , wherein the graphics processor is co-located with an application processor on a common semiconductor package. 
     
     
         15 . At least one non-transitory machine-readable medium comprising instructions that when executed by a computing device, cause the computing device to perform operations comprising:
 learning, by a graphics processor of the computing device, a structure of a generative probabilistic model;   determining, based on the structure, a stochastic inverse of the generative probabilistic model;   converting the stochastic inverse into a discriminative model;   converting the discriminative model into a deep neural network (DNN) that is trained using labeled data; and   training the DNN using labeled data, wherein training the DNN comprises at least:
 setting a bit-precision of weights are set in neurons of the DNN independently from one another; and 
 performing methodological dropout of the neurons, wherein the methodological dropout is performed in accordance with a predictivity based on historical statistical data relating to the neurons. 
   
     
     
         16 . The non-transitory machine-readable medium of  claim 15 , wherein the generative probabilistic model is unsupervised and based on unlabeled data, and wherein the discriminative model is supervised and based on labeled data, wherein the discriminative model is learned from the generative probabilistic model. 
     
     
         17 . The non-transitory machine-readable medium of  claim 15 , wherein the operations further comprise inversing the generative probabilistic model into multiple inverse models, wherein a bidirectional connection is added to connect latent variables having a common parent in each of the multiple inverse models to consolidate the multiple inverse models into a single inverse model. 
     
     
         18 . The non-transitory machine-readable medium of  claim 17 , wherein the operations further comprise converting the inverse model into the discriminative model by removing the bidirectional connection and adding a class node serving as a child node to latent leaves. 
     
     
         19 . The non-transitory machine-readable medium of  claim 15 , wherein the operations further comprise:
 generating parallel and sequential execution schedules for memory sharing at sub-network precision levels of the one or more of the neural networks; and   performing on-the-fly learning and updating of network topologies of the neural networks based on at least one of currently available data and historically available data relating to the topologies of the neural networks.   
     
     
         20 . The non-transitory machine-readable medium of  claim 15 , wherein the operations further comprise:
 facilitating at least one of an end-to-end structure learning and a sub-network structure learning; and   facilitating feature bagging or coping with large scale data by training large training sets, wherein the processor comprises a graphics processor co-located with an application processor on a common semiconductor package.   
     
     
         21 . A system comprising:
 a memory; and   a graphics processor communicably coupled to the memory, the graphics processor to:
 learn a structure of a generative probabilistic model; 
 determine, based on the structure, a stochastic inverse of the generative probabilistic model; 
 convert the stochastic inverse into a discriminative model; 
 convert the discriminative model into a deep neural network (DNN) that is trained using labeled data; and 
 train the DNN using labeled data, wherein training the DNN comprises at least:
 setting a bit-precision of weights are set in neurons of the DNN independently from one another; and 
 performing methodological dropout of the neurons, wherein the methodological dropout is performed in accordance with a predictivity based on historical statistical data relating to the neurons. 
 
   
     
     
         22 . The system of  claim 21 , wherein the generative model is unsupervised and based on unlabeled data, and wherein the discriminative model is supervised and based on labeled data, wherein the discriminative model is learned from the generative probabilistic model. 
     
     
         23 . The system of  claim 21 , wherein the graphics processor is further to inverse the generative probabilistic model into multiple inverse models, wherein a bidirectional connection is added to connect latent variables having a common parent in each of the multiple inverse models to consolidate the multiple inverse models into a single inverse model. 
     
     
         24 . The system of  claim 23 , wherein the graphics processor is further to convert the inverse model into the discriminative model by removing the bidirectional connection and adding a class node serving as a child node to latent leaves. 
     
     
         25 . The system of  claim 21 , wherein the graphics processor is further to:
 generate parallel and sequential execution schedules for memory sharing at sub-network precision levels of the one or more of the neural networks; and   perform on-the-fly learning and updating of network topologies of the neural networks based on at least one of currently available data and historically available data relating to the topologies of the neural networks.

Join the waitlist — get patent alerts

Track US2023117143A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.