US2023004536A1PendingUtilityA1

Systems and methods for a data search engine based on data profiles

Assignee: CAPITAL ONE SERVICES LLCPriority: Jul 6, 2018Filed: Sep 9, 2022Published: Jan 5, 2023
Est. expiryJul 6, 2038(~11.9 yrs left)· nominal 20-yr term from priority
G06F 16/211G06N 3/047G06N 3/045G06N 3/044G06F 16/2468G06N 5/01G06F 16/245G06N 7/01G06F 16/285G06F 16/242G06F 9/30036G06N 3/0464G06N 3/0455G06N 3/0442G06N 3/09G06N 3/0985
67
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for searching data are disclosed. For example, the system may include one or more memory units storing instructions and one or more processors configured to execute the instructions to perform operations. The operations may include receiving a sample dataset and identifying a data schema of the sample dataset. The operations may include generating a sample data vector that includes statistical metrics of the sample dataset and information based on the data schema of the sample dataset. The operations may include searching a data index comprising a plurality of stored data vectors corresponding to a plurality of reference datasets. The stored data vectors may include statistical metrics of the reference datasets and information based on corresponding data schema. The operations may include generating, based on the search and the sample data vector, one or more similarity metrics of the sample dataset to individual ones of the reference datasets.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled) 
     
     
         21 . A system for searching data, comprising:
 at least one processor; and   a memory comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:
 receiving a search request comprising a sample dataset and a vector similarity threshold of similarity between vectors; 
 in response to the received search request, performing: 
 identifying, using a data profiling model comprising a machine learning model, a data schema of the sample dataset; 
 generating a sample data vector; 
 searching a data index comprising a plurality of stored data vectors corresponding to a plurality of reference datasets, the stored data vectors comprising information describing corresponding data schemas of the reference datasets, wherein searching the data index comprises performing data schema comparisons between the data schema associated with the sample data vector and the data schemas of the reference datasets; 
 generating one or more similarity metrics of the sample dataset to individual ones of the reference datasets; 
 determining, based on the one or more similarity metrics, at least a portion of the reference datasets having at least one data vector satisfying the vector similarity threshold; and 
 returning, as a result of the received search request, the at least a portion of the reference datasets. 
   
     
     
         22 . The system of  claim 21 , the operations further comprising:
 receiving a request for at least one of a sample dataset, data vector, or data index;   retrieving the sample dataset, data vector or data index.   
     
     
         23 . The system of  claim 21 , wherein the data profiling model includes at least one of an RNN model, a CNN model, or other machine learning model. 
     
     
         24 . The system of  claim 21 , wherein the data profiling model is trained to identify complex data types. 
     
     
         25 . The system of  claim 21 , wherein the data profiling model is stored with a plurality of data profiling models in a model storage. 
     
     
         26 . The system of  claim 21 , wherein the operations further comprise performing calculations on the sample dataset. 
     
     
         27 . The system of  claim 21 , wherein the operations further comprise receiving, by an aggregator, search parameters; 
     
     
         28 . The system of  claim 27 , wherein the search parameters may include instructions to search the data index based on a comparison of data vector components of statistical metrics of the dataset and statistical metrics of variables of the dataset. 
     
     
         29 . The system of  claim 21 , wherein the operations further comprise identifying the data index. 
     
     
         30 . The system of  claim 21 , wherein the operations further comprise returning the at least a portion of the reference datasets by performing at least one of: storing the reference datasets in a data storage, providing a link to the reference datasets, or providing a compressed file of the reference datasets. 
     
     
         31 . A method for searching data, comprising:
 receiving a search request comprising a sample dataset and a vector similarity threshold of similarity between vectors;   in response to the received search request, performing:
 identifying, using a data profiling model comprising a machine learning model, a data schema of the sample dataset; 
 searching a data index comprising a plurality of stored data vectors corresponding to a plurality of reference datasets, the stored data vectors comprising information describing corresponding data schemas of the reference datasets, wherein searching the data index comprises performing data schema comparisons between the data schema associated with the sample data vector and data schemas of the reference datasets; 
 generating one or more similarity metrics of the sample dataset to individual ones of the reference datasets; 
 determining, based on the one or more similarity metrics, at least a portion of the reference datasets having at least one data vector satisfying the vector similarity threshold; and 
 returning, as a result of the received search request, the at least a portion of the reference datasets. 
   
     
     
         32 . The method of  claim 21 , the operations further comprising:
 receiving a request for at least one of a sample dataset, data vector, or data index;   retrieving the sample dataset, data vector or data index.   
     
     
         33 . The method of  claim 21 , wherein the data profiling model includes at least one of an RNN model, a CNN model, or other machine learning model. 
     
     
         34 . The method of  claim 21 , wherein the data profiling model is trained to identify complex data types. 
     
     
         35 . The method of  claim 21 , wherein the data profiling model is stored with a plurality of data profiling models in a model storage. 
     
     
         36 . The method of  claim 21 , wherein the operations further comprise performing calculations on the sample dataset. 
     
     
         37 . The method of  claim 21 , wherein the operations further comprise receiving, by an aggregator, search parameters; 
     
     
         38 . The method of  claim 37 , wherein the search parameters may include instructions to search the data index based on a comparison of data vector components of statistical metrics of the dataset and statistical metrics of variables of the dataset. 
     
     
         39 . The method of  claim 21 , wherein the operations further comprise identifying the data index. 
     
     
         40 . The method of  claim 21 , wherein the operations further comprise returning the at least a portion of the reference datasets by at least of: storing the reference datasets in a data storage, providing a link to the reference datasets, or providing a compressed file of the reference datasets.

Join the waitlist — get patent alerts

Track US2023004536A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.