Systems and methods for replacing sensitive data
Abstract
A model optimizer is disclosed for managing training of models with automatic hyperparameter tuning. The model optimizer can perform a process including multiple steps. The steps can include receiving a model generation request, retrieving from a model storage a stored model and a stored hyperparameter value for the stored model, and provisioning computing resources with the stored model according to the stored hyperparameter value to generate a first trained model. The steps can further include provisioning the computing resources with the stored model according to a new hyperparameter value to generate a second trained model, determining a satisfaction of a termination condition, storing the second trained model and the new hyperparameter value in the model storage, and providing the second trained model in response to the model generation request.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A system comprising:
at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:
receiving actual data having at least one sensitive data portion;
determining a class associated with the at least one sensitive data portion;
accessing a synthetic data generation model trained using a data space having data of the class;
generating, using the synthetic data generation model, at least one synthetic data portion; and
replacing the at least one sensitive data portion with the at least one synthetic data portion.
22 . The system of claim 21 , wherein the synthetic data generation model is trained to generate synthetic data satisfying a similarity criterion.
23 . The system of claim 22 , wherein the similarity criterion is based on at least one of a statistical correlation score, a data similarity score, or a data quality score.
24 . The system of claim 22 , wherein the synthetic data generation model is a generative adversarial network (GAN).
25 . The system of claim 21 , wherein determining the class associated with the at least one sensitive data portion comprises applying a recurrent neural network (RNN) to the actual data.
26 . The system of claim 25 , wherein the RNN is trained to distinguish between classes of data.
27 . The system of claim 21 , wherein the at least one sensitive data portion is a first text string and the at least one synthetic data portion is a second text string.
28 . The system of claim 21 , wherein the at least one sensitive data portion comprises at least one of a social security number or a financial service account number.
29 . The system of claim 21 , wherein the actual data comprises personnel records or patient medical records.
30 . The system of claim 21 , wherein the actual data comprises unstructured data, the unstructured data comprising at least one of character strings or tokens.
31 . The system of claim 21 , wherein the actual data comprises structured data, the structured data comprising at least one of a key-value pair, a relational database file, or a spreadsheet.
32 . The system of claim 21 , the operations further comprising selecting the synthetic data generation model based on the determined class.
33 . The system of claim 21 , the operations further comprising:
determining a subclass associated with the at least one sensitive data portion; and selecting the synthetic data generation model based on the determined subclass.
34 . The system of claim 33 , wherein determining the subclass comprises using a distribution model associated with the class.
35 . A method for replacing sensitive data, the method comprising:
receiving actual data having at least one sensitive data portion; determining a class associated with the at least one sensitive data portion; accessing a synthetic data generation model trained using a data space having data of the class; generating, using the synthetic data generation model, at least one synthetic data portion; and replacing the at least one sensitive data portion with the at least one synthetic data portion.
36 . The method of claim 35 , wherein the synthetic data generation model is trained to generate synthetic data satisfying a similarity criterion.
37 . The method of claim 36 , wherein the similarity criterion is based on at least one of a statistical correlation score, a data similarity score, or a data quality score.
38 . The method of claim 35 , further comprising selecting the synthetic data generation model based on the determined class.
39 . The method of claim 35 , further comprising:
determining a subclass associated with the at least one sensitive data portion; and selecting the synthetic data generation model based on the determined subclass.
40 . A non-transitory computer readable medium containing instructions that, when executed by one or more processors, cause a computing system to perform operations comprising:
receiving actual data having at least one sensitive data portion; determining, using a neural network classifier configured to distinguish classes of sensitive data, a class associated with the at least one sensitive data portion; accessing a synthetic data generation model trained using training data associated with the class; generating, using the synthetic data generation model, at least one synthetic data portion; and replacing the at least one sensitive data portion with the at least one synthetic data portion.Join the waitlist — get patent alerts
Track US2022075670A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.