US9613074B2ActiveUtilityA1

Data generation for performance evaluation

Assignee: FARAHBOD ROOZBEHPriority: Dec 23, 2013Filed: Jul 17, 2014Granted: Apr 4, 2017
Est. expiryDec 23, 2033(~7.4 yrs left)· nominal 20-yr term from priority
G06F 16/217G06F 2201/80G06F 11/3447G06F 11/3414G06F 17/30306
70
PatentIndex Score
9
Cited by
25
References
20
Claims

Abstract

The present disclosure describes methods, systems, and computer program products for generating data for performance evaluation. One computer-implemented method includes identifying a source dataset from a source database, extracting a schema defining the source database, analyzing data within the source dataset to generate a value model, the value model describing features of data in each column of each table of the source dataset, analyzing data within the source database to determine data dependency, and generating a data specification file combining the extracted schema, the value model, and the data dependencies.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A computer-implemented method executed by one or more processors, the method comprising:
 identifying a source dataset from a source database; 
 extracting a schema defining the source database; 
 analyzing data within the source dataset to generate a value model, the value model describing features of data in each column of each table of the source dataset; 
 analyzing data within the source database to determine data dependency; and 
 generating a data specification file combining the extracted schema, the value model, and the data dependencies. 
 
     
     
       2. The method of  claim 1 , wherein determining data dependency comprises determining one or more of inter-table dependency, intra-table dependency, or inter-column dependency of data within the source database. 
     
     
       3. The method of  claim 1 , comprising storing the data specification file in a database or writing the data specification file in a file accessible by a user. 
     
     
       4. The method of  claim 1 , comprising generating synthetic data based on the data specification file, wherein the generated synthetic data corresponds to the extracted schema of the source database, wherein the generated synthetic data corresponds to the generated value model of the source database, wherein the generated synthetic data corresponds to the determined data dependency, and wherein the generated synthetic data includes different data than the analyzed data within the source database. 
     
     
       5. The method of  claim 4 , wherein the generated synthetic data has a smaller size, a same size, or a larger size than the analyzed data. 
     
     
       6. The method of  claim 4 , wherein generating synthetic data comprises:
 identifying a data source; 
 identifying a value model for the synthetic data based on the generated value model of the source database; 
 generating a composite value model based on the determined data dependency; and 
 generating row values of the synthetic data based on the data source, the value model, and the composite value model. 
 
     
     
       7. The method of  claim 4 , wherein generating synthetic data comprises:
 generating a foreign key graph; and 
 generating a data tables in an order according to the foreign key graph. 
 
     
     
       8. A non-transitory, computer-readable medium storing computer-readable instructions executable by a computer and configured to:
 identify a source dataset from a source database; 
 extract a schema defining the source database; 
 analyze data within the source dataset to generate a value model, the value model describing features of data in each column of each table of the source dataset; 
 analyze data within the source database to determine data dependency; and 
 generate a data specification file combining the extracted schema, the value model, and the data dependencies. 
 
     
     
       9. The medium of  claim 8 , wherein determining data dependency comprises determining one or more of inter-table dependency, intra-table dependency, or inter-column dependency of data within the source database. 
     
     
       10. The medium of  claim 8 , comprising instructions to store the data specification file in a database or write the data specification file in a file accessible by a user. 
     
     
       11. The medium of  claim 8 , comprising instructions to generate synthetic data based on the data specification file, wherein the generated synthetic data corresponds to the extracted schema of the source database, wherein the generated synthetic data corresponds to the generated value model of the source database, wherein the generated synthetic data corresponds to the determined data dependency, and wherein the generated synthetic data includes different data than the analyzed data within the source database. 
     
     
       12. The medium of  claim 11 , wherein the generated synthetic data has a smaller size, a same size, or a larger size than the analyzed data. 
     
     
       13. The medium of  claim 11 , wherein generating synthetic data comprises:
 identifying a data source; 
 identifying a value model for the synthetic data based on the generated value model of the source database; 
 generating a composite value model based on the determined data dependency; and 
 generating row values of the synthetic data based on the data source, the value model, and the composite value model. 
 
     
     
       14. The medium of  claim 11 , wherein generating synthetic data comprises:
 generating a foreign key graph; and 
 generating a data tables in an order according to the foreign key graph. 
 
     
     
       15. A system, comprising:
 a memory; 
 at least one hardware processor interoperably coupled with the memory and configured to:
 identify a source dataset from a source database; 
 extract a schema defining the source database; 
 analyze data within the source dataset to generate a value model, the value model describing features of data in each column of each table of the source dataset; 
 analyze data within the source database to determine data dependency; and 
 generate a data specification file combining the extracted schema, the value model, and the data dependencies. 
 
 
     
     
       16. The system of  claim 15 , wherein determining data dependency comprises determining one or more of inter-table dependency, intra-table dependency, or inter-column dependency of data within the source database. 
     
     
       17. The system of  claim 15 , comprising instructions to store the data specification file in a database or write the data specification file in a file accessible by a user. 
     
     
       18. The system of  claim 15 , comprising instructions to generate synthetic data based on the data specification file, wherein the generated synthetic data corresponds to the extracted schema of the source database, wherein the generated synthetic data corresponds to the generated value model of the source database, wherein the generated synthetic data corresponds to the determined data dependency, and wherein the generated synthetic data includes different data than the analyzed data within the source database. 
     
     
       19. The system of  claim 18 , wherein the generated synthetic data has a smaller size, a same size, or a larger size than the analyzed data. 
     
     
       20. The system of  claim 18 , wherein generating synthetic data comprises:
 identifying a data source; 
 identifying a value model for the synthetic data based on the generated value model of the source database; 
 generating a composite value model based on the determined data dependency; 
 generating a foreign key graph; and 
 generating row values of the synthetic data based on the data source, the value model, the composite value model and the foreign key graph.

Join the waitlist — get patent alerts

Track US9613074B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.