Data set and algorithm validation, bias characterization, and valuation
Abstract
A system and method for providing access and agency to individual entities and people over their data for the purpose of data set validation to facilitate data set and algorithm bias certification and scoring. A first data set is filtered to extract its core information content and to create a certified data set. A certified model is created by training a machine learning algorithm on the certified data set, which certified model is then used to evaluate the bias of subsequent data sets. The data set may be given a value score which represents the overall validity of the data set and its bias characterization. A bias characterization audit can help identify the root causes of bias outcomes from predictive software and algorithms that perform third party tasks and services. The score can be used as a metric to further facilitate market transactions.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system for data set validation, bias characterization, and valuation, comprising:
a computing device comprising a memory, a processor, and a non-volatile data storage device; a data set and model certification manager comprising a first plurality of programming instructions stored in the memory of, and operating on the processor of, the computing device, wherein the first plurality of programming instructions, when operating on the processor, cause the computing device to:
retrieve a first data set from the non-volatile data storage device;
pass the first data set through a series of filters to reduce the first data set to its core information content;
analyze the core information content to determine an information gain for the first data set based on an entropy of the core information content;
certify the data set if the information gain exceeds a threshold;
create a certified model by training a machine learning algorithm with the certified data set;
use the certified model to generate a baseline output using the first data set as input; and
store the certified data set, the certified model, and the baseline output in the non-volatile data storage device; and
a bias characterization auditor comprising a second plurality of programming instructions stored in the memory of, and operating on the processor of, the computing device, wherein the second plurality of programming instructions, when operating on the processor, cause the computing device to:
receive a second data set;
retrieve the certified model and the baseline output from the non-volatile data storage device;
use the second data set as an input to the certified model to generate a set output;
perform a bias characterization analysis by comparing the baseline output to the set output;
generate a bias characterization score from the bias characterization analysis; and
store the bias characterization score in the non-volatile data storage device; and
a data valuation engine comprising a third plurality of programming instructions stored in the memory of, and operating on the processor of, the computing device, wherein the third plurality of programming instructions, when operating on the processor, cause the computing device to:
score the second data set based on a plurality of scoring metrics, one of which is the bias characterization score; and
create and store a data set value score as a weighted combination of the scores of the plurality of scoring metrics.
2 . The system of claim 1 , wherein the value score is used as a pricing schedule for data set monetization.
3 . The system of claim 1 , wherein the bias characterization auditor further receives a bias audit claim containing an audited data set and performs the bias characterization analysis on the audited data set to create a bias characterization score for the audited data set.
4 . The system of claim 1 , wherein the second data set consists of: partial data, statistical characteristic data, synthetic data, a model characterizing synthetic data, or tokenized data.
5 . The system of claim 1 , further comprising a data translator comprising a fourth plurality of programming instructions stored in the memory and operating on the processor which cause the computing device to translate a data set into one or more optional data representations while maintaining links to the original source or sources of data.
6 . A method for data set validation and valuation to facilitate data set and algorithm bias certification and scoring, comprising the steps of:
retrieving a first data set; passing the first data set through a series of filters to reduce the first data set to its core information content; analyzing the core information content to determine an information gain for the first data set based on an entropy of the core information content; certifying the first data set if the information gain exceeds a threshold; creating a certified model by training a machine learning algorithm with the certified data set; using the certified model to generate a baseline output using the first data set as input; storing the certified data set, certified model, and the baseline output; receiving a second data set using the second data set as an input to the certified model to generate a set output; performing a bias characterization analysis by comparing the baseline output to the set output; generating a bias characterization score from the bias characterization analysis; storing the bias characterization score; scoring the second data set based on a plurality of scoring metrics, one of which is the bias characterization score; and creating and storing a data set value score as a weighted combination of the scores of the plurality of scoring metrics.
7 . The method of claim 6 , wherein the value score is used as a pricing schedule for data set monetization.
8 . The method of claim 6 , further comprising the steps of receiving a bias audit claim containing an audited data set and performing the bias characterization analysis on the audited data set to create a bias characterization score for the audited data set.
9 . The method of claim 6 , wherein the second data set consists of: partial data, statistical characteristic data, synthetic data, a model characterizing synthetic data, or tokenized data.
10 . The method of claim 6 , further comprising a data translator further comprising the steps of translating a data set into one or more optional data representations while maintaining links to the original source or sources of data.Join the waitlist — get patent alerts
Track US2021112101A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.