Systems and methods for synthetic data generation
Abstract
A cloud computing system can be configured to generate data models. A model optimizer of the cloud computing system can provision computing resources of the cloud computing system with a data model. A dataset generator of the cloud computing system can generate a synthetic dataset for training the data model. The computing resources can train the data model using the synthetic dataset. The model optimizer can store the data model and metadata of the data model in a model storage. The cloud computing system can receive production data from a data source by a production instance of the cloud computing system using a common file system. The production data can be processed using the data model by the production instance. The computing resources, the dataset generator, and the model optimizer can be hosted by separate virtual computing instances of the cloud computing system.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A cloud computing system for generating data models, comprising:
at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor cause the cloud computing system to perform operations comprising:
normalizing a reference dataset;
receiving a similarity criterion, the similarity criterion including a predetermined difference in value between the normalized reference dataset and an output dataset of a data model;
generating a synthetic dataset for training the data model;
training the data model using the synthetic dataset, the training comprising:
generating, based on a comparison of the output dataset and the normalized reference dataset, a similarity metric of the data model,
generating a prediction metric of the data model,
evaluating the similarity metric against the similarity criterion,
evaluating the prediction metric against a prediction criterion, and
updating the data model based on the evaluations of the similarity metric and prediction metric, the updating comprising penalizing generation of synthetic data by adding a penalty term to a loss function;
repeating the training until the similarity criterion is met by the similarity metric and the prediction criterion is met by the similarity metric and the prediction metric; and
in response to the similarity criterion being met by the similarity metric and the prediction criterion being met by the prediction metric, storing the data model in a model storage.
22 . The cloud computing system of claim 21 , wherein the similarity metric depends on a maximum distance or an average distance according to a distance measure between rows selected from the output dataset and at least one row selected from the reference dataset.
23 . The cloud computing system of claim 21 , wherein the loss function is updatable for training the data model, the loss function is associated with a penalty term, to ensure that the value of the similarity metric exceeds a similarity threshold or remains near the similarity threshold.
24 . The cloud computing system of claim 21 , wherein:
the similarity metric comprises at least one of a statistical correlation score, data similarity score, or data quality score; and the prediction metric includes at least one of a prediction accuracy verification, a prediction accuracy cross validation, a regression verification, a regression cross validation, or a principal component analysis.
25 . The cloud computing system of claim 24 , wherein the similarity metric is configured to calculate scores using the synthetic dataset and a reference dataset.
26 . The cloud computing system of claim 21 , wherein the synthetic dataset differs in value from the normalized reference dataset according to a predetermined amount according to the similarity metric.
27 . The cloud computing system of claim 21 , wherein:
the similarity metric depends on a covariance of the synthetic dataset and a covariance of the normalized reference dataset; and the operations further comprise generating a difference matrix using a covariance matrix of the normalized reference dataset and a covariance matrix of the synthetic dataset.
28 . The cloud computing system of claim 21 , wherein the prediction metric includes at least one of a prediction accuracy check, a prediction accuracy cross check, a regression check, a regression cross check, or a principal component analysis check.
29 . The cloud computing system of claim 21 , wherein the similarity metric depends on one or more criteria, the one or more criteria comprising at least one of:
a covariance of output dataset and a covariance of the normalized reference dataset; a univariate value distribution of an element of the synthetic dataset; a univariate value distribution of an element of the normalized reference dataset; a number of elements of the synthetic dataset that match elements of the reference dataset; a number of elements of the synthetic dataset that are similar to elements of the normalized reference dataset; a distance measure between each row of the synthetic dataset and each row of the normalized reference dataset; a frequency of duplicate elements in the synthetic dataset and the normalized reference dataset; and a relative prevalence of rare values in the synthetic dataset and the normalized reference dataset; and differences in ratios between the synthetic dataset and the normalized reference dataset.
30 . The cloud computing system of claim 21 , wherein the similarity criterion concerns at least one of a statistical correlation score between the synthetic data and the normalized reference dataset, a data similarity score between the synthetic dataset and the reference dataset, or a data quality score for the synthetic dataset.
31 . A method for generating data models, comprising:
normalizing a reference dataset; receiving a similarity criterion, the similarity criterion including a predetermined difference in value between the normalized reference dataset and an output dataset of the data model; generating a synthetic dataset for training the data model; training a data model using the synthetic dataset, the training comprising:
generating, based on a comparison of the output dataset and the normalized reference dataset, a similarity metric of the data model,
generating a prediction metric of the data model,
evaluating the similarity metric against the similarity criterion,
evaluating the prediction metric against a prediction criterion, and
updating the data model based on the evaluations of the similarity metric and prediction metric, the updating comprising penalizing generation of synthetic data by adding a penalty term to a loss function;
repeating the training until the similarity criterion is met by the similarity metric and the prediction criterion is met by the similarity metric and the prediction metric; and in response to the similarity criterion being met by the similarity metric and the prediction criterion being met by the prediction metric, storing the data model and new metadata in a model storage.
32 . The method of claim 31 , wherein the similarity criterion concerns at least one of a statistical correlation score between the synthetic data and the normalized reference dataset, a data similarity score between the synthetic dataset and the reference dataset, or a data quality score for the synthetic dataset.
33 . The method of claim 31 , wherein the similarity metric depends on a maximum distance or an average distance according to a distance measure between rows selected from the output dataset and at least one row selected from the reference dataset.
34 . The method of claim 31 , wherein the similarity metric comprises at least one of a statistical correlation score, data similarity score, or data quality score, and the prediction metric includes at least one of a prediction accuracy verification, a prediction accuracy cross validation, a regression verification, a regression cross validation, or a principal component analysis.
35 . The method of claim 34 , wherein the similarity metric is configured to calculate scores using the synthetic dataset and a reference dataset.
36 . The method of claim 31 , wherein the synthetic dataset differs in value from the normalized reference dataset according to a predetermined amount according to the similarity metric.
37 . The method of claim 31 , wherein:
the similarity metric depends on a covariance of the synthetic dataset and a covariance of the normalized reference dataset; and the operations further comprise generating a difference matrix using a covariance matrix of the normalized reference dataset and a covariance matrix of the synthetic dataset.
38 . The method of claim 31 , wherein the prediction metric includes at least one of a prediction accuracy check, a prediction accuracy cross check, a regression check, a regression cross check, or a principal component analysis check.
39 . The method of claim 31 , wherein the metadata includes an indication of origin of the new synthetic data model and the data used to generate the new synthetic data model.
40 . A non-transitory computer-readable memory storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:
normalizing a reference dataset; receiving a similarity criterion, the similarity criterion including a predetermined difference in value between the normalized reference dataset and an output dataset of the data model; generating a synthetic dataset for training the data model; training a data model using the synthetic dataset, the training comprising:
generating, based on a comparison of the output dataset and the normalized reference dataset, a similarity metric of the data model,
generating a prediction metric of the data model,
evaluating the similarity metric against the similarity criterion,
evaluating the prediction metric against a prediction criterion to determine whether data models perform similarly for both the synthetic data and actual data, and
updating the data model based on the evaluations of the similarity metric and prediction metric, the updating comprising penalizing generation of synthetic data by adding a penalty term to a loss function, decreasing values of the similarity metric to indicate similarity;
repeating the training until the similarity criterion is met by the similarity metric and the prediction criterion is met by the similarity metric and the prediction metric; and in response to the similarity criterion being met by the similarity metric and the prediction criterion being met by the prediction metric, storing the data model in a model storage, once a value of a similarity metric or prediction metric satisfies a predetermined threshold.Join the waitlist — get patent alerts
Track US2023195541A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.