US2021092160A1PendingUtilityA1

Data set creation with crowd-based reinforcement

Assignee: QOMPLX INCPriority: Oct 28, 2015Filed: Aug 3, 2020Published: Mar 25, 2021
Est. expiryOct 28, 2035(~9.2 yrs left)· nominal 20-yr term from priority
G06N 3/045G06N 3/094G06N 3/0475H04L 63/1425G06F 16/951G06F 16/2477H04L 63/1441H04L 63/20
63
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method for creation and expansion of high quality data set collections for training of machine learning algorithms via crowdsourced curation that utilizes a data marketplace which incentivizes data gatherers, publishers, and users to contribute to the creation of a vast resource of reliable data set collections. Data is automatically ingested from disparate sources and autonomously checked for data quality, provenance, and cyber-risks and subsequently given a reputation score. Data stewards curate a queue of low scoring real data as well as synthetically generated data. All reputable data is stored for user consumption and further iterative data generation.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system for creation and expansion of high quality data set collections for training of machine learning algorithms via crowdsourced curation, comprising:
 a reputation scoring engine comprising a first plurality of programming instructions stored in a memory of, and operating on a processor of, a computing device, wherein the first plurality of programming instructions, when operating on the processor, cause the computing device to:
 receive a data set; 
 score a data entry within the data set, wherein the score is calculated from a plurality of scoring metrics; 
 sum all of the data entry scores within the data set combining to form an overall reputation score; 
 flag an erroneous data entry which may not be resolved through the machine learning algorithm; 
 compare the overall reputation score with a numerical threshold for reputability; 
 send the flagged erroneous data entry and the data sets not meeting the threshold for reputability to a verification queue; and 
 store the data sets that meet the threshold for reputability to a data store as a reputable data set collection; and 
   a verification queue comprising a second plurality of programming instructions stored in the memory of, and operating on the processor of, the computing device, wherein the second plurality of programming instructions, when operating on the processor, cause the computing device to:
 receive the flagged erroneous data entry and the data sets not meeting the threshold for reputability; 
 assign the data in the verification queue to a data steward for human curation; and 
 send the curated and resolved data back to the reputation scoring engine for an additional iteration. 
   
     
     
         2 . The system of  claim 1 , further comprising a synthetic data generator comprising a third plurality of programming instructions stored in the memory of, and operating on the processor of, the computing device, wherein the third plurality of programming instructions, when operating on the processor, cause the computing device to:
 retrieve a reputable data set collection stored within the data store;   generate a synthetic data set from the reputable data set;   send the synthetic data set to the reputation scoring engine; and   wherein the reputation scoring engine further merges the synthetic data with the reputable data set collection where the synthetic data set passes the threshold for reputability.   
     
     
         3 . The system of  claim 2 , wherein the synthetic data generator is a generative adversarial network. 
     
     
         4 . A method for creation and expansion of high quality data set collections for training of machine learning algorithms via crowdsourced curation, comprising the steps of:
 receiving a data set;
 scoring a data entry within the data set, wherein the score is calculated from a plurality of scoring metrics; 
 summing all of the data entry scores within the data set combining to form an overall reputation score; 
 flagging an erroneous data entry which may not be resolved through the machine learning algorithm; 
 comparing the overall reputation score with a numerical threshold for reputability; 
 sending the flagged erroneous data entry and the data sets not meeting the threshold for reputability to a verification queue; 
 storing the data sets that meet the threshold for reputability to a data store; 
 assigning the data in the verification queue to a data steward for human curation; and 
 sending the curated data back for an additional iteration of the above process. 
   
     
     
         5 . The method of  claim 4 , further comprising the steps of:
 retrieving a data set collection stored within the data store;   generating a synthetic data set from the retrieved data set;   sending the synthetic data set for scoring based on the above process; and   merging the synthetic data with the reputable data set collection where the synthetic data set passes the threshold for reputability.   
     
     
         6 . The method of  claim 5 , wherein the synthetic data set is generated by a generative adversarial network.

Join the waitlist — get patent alerts

Track US2021092160A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.