Data set creation with crowd-based reinforcement
Abstract
A system and method for creation and expansion of high quality data set collections for training of machine learning algorithms via crowdsourced curation that utilizes a data marketplace which incentivizes data gatherers, publishers, and users to contribute to the creation of a vast resource of reliable data set collections. Data is automatically ingested from disparate sources and autonomously checked for data quality, provenance, and cyber-risks and subsequently given a reputation score. Data stewards curate a queue of low scoring real data as well as synthetically generated data. All reputable data is stored for user consumption and further iterative data generation.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system for creation and expansion of high quality data set collections for training of machine learning algorithms via crowdsourced curation, comprising:
a reputation scoring engine comprising a first plurality of programming instructions stored in a memory of, and operating on a processor of, a computing device, wherein the first plurality of programming instructions, when operating on the processor, cause the computing device to:
receive a data set;
score a data entry within the data set, wherein the score is calculated from a plurality of scoring metrics;
sum all of the data entry scores within the data set combining to form an overall reputation score;
flag an erroneous data entry which may not be resolved through the machine learning algorithm;
compare the overall reputation score with a numerical threshold for reputability;
send the flagged erroneous data entry and the data sets not meeting the threshold for reputability to a verification queue; and
store the data sets that meet the threshold for reputability to a data store as a reputable data set collection; and
a verification queue comprising a second plurality of programming instructions stored in the memory of, and operating on the processor of, the computing device, wherein the second plurality of programming instructions, when operating on the processor, cause the computing device to:
receive the flagged erroneous data entry and the data sets not meeting the threshold for reputability;
assign the data in the verification queue to a data steward for human curation; and
send the curated and resolved data back to the reputation scoring engine for an additional iteration.
2 . The system of claim 1 , further comprising a synthetic data generator comprising a third plurality of programming instructions stored in the memory of, and operating on the processor of, the computing device, wherein the third plurality of programming instructions, when operating on the processor, cause the computing device to:
retrieve a reputable data set collection stored within the data store; generate a synthetic data set from the reputable data set; send the synthetic data set to the reputation scoring engine; and wherein the reputation scoring engine further merges the synthetic data with the reputable data set collection where the synthetic data set passes the threshold for reputability.
3 . The system of claim 2 , wherein the synthetic data generator is a generative adversarial network.
4 . A method for creation and expansion of high quality data set collections for training of machine learning algorithms via crowdsourced curation, comprising the steps of:
receiving a data set;
scoring a data entry within the data set, wherein the score is calculated from a plurality of scoring metrics;
summing all of the data entry scores within the data set combining to form an overall reputation score;
flagging an erroneous data entry which may not be resolved through the machine learning algorithm;
comparing the overall reputation score with a numerical threshold for reputability;
sending the flagged erroneous data entry and the data sets not meeting the threshold for reputability to a verification queue;
storing the data sets that meet the threshold for reputability to a data store;
assigning the data in the verification queue to a data steward for human curation; and
sending the curated data back for an additional iteration of the above process.
5 . The method of claim 4 , further comprising the steps of:
retrieving a data set collection stored within the data store; generating a synthetic data set from the retrieved data set; sending the synthetic data set for scoring based on the above process; and merging the synthetic data with the reputable data set collection where the synthetic data set passes the threshold for reputability.
6 . The method of claim 5 , wherein the synthetic data set is generated by a generative adversarial network.Join the waitlist — get patent alerts
Track US2021092160A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.