US2020090050A1PendingUtilityA1

Systems and methods for training generative machine learning models with sparse latent spaces

Assignee: D WAVE SYSTEMS INCPriority: Sep 14, 2018Filed: Sep 5, 2019Published: Mar 19, 2020
Est. expirySep 14, 2038(~12.1 yrs left)· nominal 20-yr term from priority
G06N 3/088G06N 3/0445G06N 3/047G06N 20/00G06N 7/01G06N 3/044G06N 3/0495G06N 3/082G06N 3/0455G06N 3/0464G06N 3/0475
46
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Generative machine learning models, such as variational autoencoders, with comparatively sparse latent spaces are provided. Continuous latent variables are activated and/or inactivated based on a state of the latent space. Activation may be controlled by corresponding binary latent variables and/or by rectification of probability distributions defined over the latent space. Sparsification may be supported by normalization of terms, such as providing an L1 or L2 prior.

Claims

exact text as granted — not AI-modified
1 . A method for unsupervised learning over an input space comprising a plurality of input variables, and at least a subset of a training dataset of samples of the respective variables, to attempt to identify the value of at least one parameter that increases the log-likelihood of the at least a subset of a training dataset with respect to a model, the model expressible as a function of the at least one parameter, the method executed by circuitry including at least one processor and comprising;
 forming a latent space comprising a plurality of continuous random latent variables;   forming an approximating posterior distribution over the latent space, conditioned on the input space, and formed by, for each of the continuous random latent variables, truncating a corresponding encoding base distribution based on input data from the input space;   forming a prior distribution over the latent space;   forming a decoding distribution over the input space; and   training the model based on the encoding, prior, and decoding distributions.   
     
     
         2 . The method of  claim 1  wherein forming the prior distribution comprises, for each of the continuous random latent variables, truncating a corresponding prior base distribution by rectifying the corresponding prior base distribution based on the continuous random latent variable. 
     
     
         3 . The method of  claim 2  wherein, for each continuous random latent variable, the corresponding encoding base distribution and the corresponding prior base distribution are parametrizations of a shared distribution, forming the prior distribution comprises truncating the shared distribution, and forming the approximating posterior distribution comprises truncating the shared distribution. 
     
     
         4 . The method of  claim 3  wherein the shared distribution comprises a Gaussian distribution and truncating the shared distribution comprises truncating the Gaussian distribution. 
     
     
         5 . The method of  claim 1  wherein, when forming the approximating posterior, truncating the corresponding encoding base distribution comprises rectifying at least one of the continuous random latent variables. 
     
     
         6 . The method of  claim 5  wherein training the model comprises determining a gradient over the approximating posterior based on a reparametrization of the at least one of the continuous random latent variables. 
     
     
         7 . The method of  claim 5  wherein rectifying at least one of the continuous random latent variables comprises applying a rectified linear unit to an initial value of the at least one of the continuous random latent variables generated by the approximating posterior distribution. 
     
     
         8 . The method of  claim 1  wherein forming the latent space further comprises forming a plurality of discrete random latent variables and, for each of the plurality of continuous variables, truncating the corresponding prior base distribution comprises truncating the corresponding prior base distribution based on a state of a corresponding one of the discrete random latent variables. 
     
     
         9 . The method of  claim 8  wherein, for each of the plurality of continuous variables, truncating the corresponding prior base distribution based on the state of the corresponding one of the discrete random latent variables comprises selecting at least one of: an activation regime and an inactivation regime and:
 if the activation regime is selected, causing samples to be drawn for the continuous random variable from the corresponding prior base distribution; and 
 if the inactivation regime is selected, causing samples to be drawn for the continuous random variable from a singularity distribution. 
 
     
     
         10 . The method of  claim 9  wherein the singularity distribution comprises a Dirac delta distribution. 
     
     
         11 . The method of  claim 9  wherein training the model comprises regularizing one or more continuous random latent variables based on the one or more continuous random latent variables being in the activation regime. 
     
     
         12 . The method of  claim 1  wherein each of a first subset of the plurality of continuous random latent variables share a first common base distribution and forming the approximating posterior distribution comprises, for each of the first subset, truncating a corresponding approximating posterior base distribution comprises truncating the first common base distribution. 
     
     
         13 . The method of  claim 12  wherein training the model comprises determining a gradient of an objective function based on a reparametrization of the first subset of continuous random latent variables. 
     
     
         14 . The method of  claim 13  wherein:
 each of a second subset of the plurality of continuous random latent variables share a second common base distribution, the second common base distribution having at least one trainable parameter separate from the one or more trainable parameters of the first common base distribution; and 
 forming the approximating posterior distribution comprises, for each continuous random latent variable of the second subset, truncating a corresponding approximating posterior base distribution comprises truncating the first common base distribution. 
 
     
     
         15 .- 26 . (canceled) 
     
     
         27 . A computational system, comprising:
 at least one processor; and   at least one nontransitory processor-readable storage medium that stores at least one of processor-executable instructions or data which, when executed by the at least one processor cause the at least one processor to:
 form a latent space comprising a plurality of continuous random latent variables; 
 form an approximating posterior distribution over the latent space, conditioned on the input space, and formed by, for each of the continuous random latent variables, truncating a corresponding encoding base distribution based on input data from the input space; 
 form a prior distribution over the latent space; 
 form a decoding distribution over the input space; and 
 train the model based on the encoding, prior, and decoding distributions.

Join the waitlist — get patent alerts

Track US2020090050A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.