Discrete variational auto-encoder systems and methods for machine learning using adiabatic quantum computers
Abstract
A computational system can include digital circuitry and analog circuitry, for instance a digital processor and a quantum processor. The quantum processor can operate as a sample generator providing samples. Samples can be employed by the digital processing in implementing various machine learning techniques. For example, the computational system can perform unsupervised learning over an input space, for example via a discrete variational auto-encoder, and attempting to maximize the log-likelihood of an observed dataset. Maximizing the log-likelihood of the observed dataset can include generating a hierarchical approximating posterior. Unsupervised learning can include generating samples of a prior distribution using the quantum processor. Generating samples using the quantum processor can include forming chains of qubits and representing discrete variables by chains.
Claims
exact text as granted — not AI-modified1 . A method for searching an input space characterized by an objective function and comprising discrete or continuous variables, to attempt to identify an element of the input space based on the objective function, the method executed by circuitry including at least one processor, the method comprising:
training a machine learning model based on at least a subset of a training dataset of samples of respective variables of the input space, said training comprising:
forming a latent representation of a latent space based on the input space and comprising a plurality of random variables, the plurality of random variables comprising one or more discrete random variables; and
forming an encoding distribution over the latent space, conditioned on the input space;
searching the latent space based on at least a subset of the one or more discrete random variables and the objective function, said searching comprising:
selecting an initial point in the latent space based on the latent representation;
determining one or more objective values for one or more points in the input space based on the objective function;
determining one or more points in the latent space corresponding to the input space based on the encoding distribution; and
optimizing the initial point in the latent space based on the latent representation and at least a subset of the one or more discrete random variables, said optimizing comprising associating the one or more objective values with the one or more points in the latent space.
2 . The method of claim 1 wherein:
training further comprises forming a decoding distribution over the input space conditioned on the latent representation; and
associating the one or more objective values with the one or more points in the latent space comprises determining a composition of the objective function with the decoding distribution to provide a composed objective function with a domain in the latent space.
3 . The method of claim 2 wherein training a machine learning model comprises training a variational autoencoder, forming an encoding distribution comprises forming an approximating posterior distribution, and forming a decoding distribution comprises forming a conditional distribution over the input space.
4 . The method of claim 1 wherein forming a latent representation of a latent space based on the input space and comprising a plurality of random variables comprises forming a latent representation of a latent space comprising one or more continuous random variables and wherein optimizing the initial point in the latent space comprises optimizing over at least a subset of the one or more continuous random variables.
5 . The method of claim 4 wherein optimizing comprises one of fixing and varying at least one of the one or more discrete random variables and varying at least one of the one or more continuous random variables.
6 . The method of claim 1 wherein searching comprises selecting a plurality of initial points and optimizing the plurality of initials point in the latent space based on the latent representation and at least a subset of the one or more discrete random variables.
7 . The method of claim 1 wherein searching comprises determining a result in the latent space and decoding the result into the input space via a decoding distribution over the input space conditioned on the latent representation.
8 . The method of claim 1 wherein training the machine learning model based on at least a subset of a training dataset of samples comprises training the machine learning model based on at least a subset of a training dataset of samples comprising supervised data and forming the latent representation of the latent space based on the input space comprises forming a latent representation of a latent space based on at least a subset of the supervised data.
9 . The method of claim 8 wherein searching comprises optimizing the latent representation based on a gradient over the latent space, the gradient based on the objective function, the objective function comprising a loss function defined on one or more properties represented in the supervised data.
10 . The method of claim 9 wherein optimizing the latent representation based on the gradient comprises backpropagating values of the loss function from the input space to the latent space via a decoding distribution over the input space conditioned on the latent representation.
11 . The method of claim 10 wherein optimizing comprises optimizing based on both a log-probability of the properties and a log probability of a prior distribution over the latent space.
12 . A computational system, comprising:
at least one processor; and at least one nontransitory processor-readable storage medium that stores at least one of processor-executable instructions or data which, when executed by the at least one processor cause the at least one processor to:
train a machine learning model, based on at least a subset of a training dataset of samples of respective variables of an input space, the input space characterized by an objective function and comprising discrete or continuous variables, to:
form a latent representation of a latent space based on the input space and comprising a plurality of random variables, the plurality of random variables comprising one or more discrete random variables; and
form an encoding distribution over the latent space, conditioned on the input space;
search the latent space, based on at least a subset of the one or more discrete random variables and the objective function, to:
select an initial point in the latent space based on the latent representation;
determine one or more objective values for one or more points in the input space based on the objective function;
determine one or more points in the latent space corresponding to the input space based on the encoding distribution; and
optimize the initial point in the latent space, based on the latent representation and at least a subset of the one or more discrete random variables, to associate the one or more objective values with the one or more points in the latent space.
13 . The computational system of claim 12 wherein:
the machine learning model is trained to form a decoding distribution over the input space conditioned on the latent representation; and
the one or more objective values associate with the one or more points in the latent space to determine a composition of the objective function with the decoding distribution to provide a composed objective function with a domain in the latent space.
14 . The computational system of claim 13 wherein the machine learning model comprises a variational autoencoder, the encoding distribution comprises an approximating posterior distribution, and the decoding distribution comprises a conditional distribution over the input space.
15 . The computational system of claim 12 wherein the plurality of random variables comprises one or more continuous random variables and the initial point in the latent space is optimized over at least a subset of the one or more continuous random variables.
16 . The computational system of claim 15 wherein the at least one of processor-executable instructions or data cause the at least one processor to optimize the initial point in the latent space comprising at least one of the one or more discrete random variables being one of fixed and varied and at least one of the one or more continuous random variables being varied.
17 . The computational system of claim 12 wherein the search of the latent space comprises a selection of a plurality of initial points and an optimization of the plurality of initial point in the latent space based on the latent representation and at least a subset of the one or more discrete random variables.
18 . The computational system of claim 12 wherein the search of the latent space comprises a determination of a result in the latent space and the result being decoded into the input space via a decoding distribution over the input space conditioned on the latent representation.
19 . The computational system of claim 12 wherein the training dataset comprises supervised data and the latent representation of the latent space is formed based on at least a subset of the supervised data.
20 . The computational system of claim 19 wherein:
the objective function comprises a loss function defined on one or more properties represented in the supervised data; and
the latent representation is optimized based on a gradient over the latent space, the gradient based on the loss function.
21 . The computational system of claim 20 wherein the latent representation is optimized by backpropagating values of the loss function from the input space to the latent space via a decoding distribution over the input space conditioned on the latent representation.
22 . The computational system of claim 20 wherein the loss function is defined only on the one or more properties represented in the supervised data.
23 . The computational system of claim 20 wherein the latent representation is optimized based on both a log-probability of the properties and a log probability of a prior distribution over the latent space.Join the waitlist — get patent alerts
Track US2021365826A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.