US2016078367A1PendingUtilityA1
Data clean-up method for improving predictive model training
Est. expiryOct 15, 2034(~8.2 yrs left)· nominal 20-yr term from priority
Inventors:Akli Adjaoute
G06F 16/215G06N 5/04G06N 99/005G06N 20/00
56
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A method that improves the training of predictive models. Better trained predictive models make better predictions, and can classify transactions with reduced levels of false positives and false negative. Included is an apparatus for executing a data clean-up algorithm that harmonizes a wide range of real world supervised and unsupervised training data into a single, error-free, uniformly formatted record file that has every field coherent and well populated with information.
Claims
exact text as granted — not AI-modified1 . A method that improves the training of predictive models, comprising:
converting and transforming a variety of inconsistent and incoherent supervised and unsupervised training data for predictive models received by a network server as electronic data files, and storing that in a computer data storage mechanism, and then into another single, error-free, uniformly formatted record file stored in the computer data storage mechanism with an apparatus for executing a data integrity analysis algorithm that harmonizes a range of supervised and unsupervised training data into flat-data records in which every field of every record file is modified to be coherent and well-populated with information; comparing and correcting any data values in each data field in the inconsistent and incoherent supervised and unsupervised training data according to a user-service consumer preference and a predefined data dictionary of valid data values with an apparatus for executing an algorithm that substitutes data values in the data fields of incoming supervised and unsupervised training data with at least one value representing a minimum, a maximum, a null, an average, and a default; discerning the context of any text included in the inconsistent and incoherent supervised and unsupervised training data with an apparatus for executing a contextual dictionary algorithm that employs a thesaurus of alternative contexts of ambiguous words for find a common context denominator, and to then record the context determined into the computer data storage mechanism for later access by a predictive model; cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for executing an algorithm for cleaning up raw data in stored data records, field-by-field, record-by-record in which some types of fields are restricted in what is legal or allowed, and includes fetching raw data from the computer data storage mechanism and testing each field if a data value reported is numeric or symbolic, and if numeric, a data dictionary is used to see if such data value is previously listed as valid, and if symbolic, using another data dictionary to see if such data value is listed there as valid; cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for executing a Smith-Waterman algorithm for a local-sequence alignment and to determine if there are any similar regions between two strings or sequences, and in which a consistent, coherent terminology is then enforceable in each data field without data loss, and in which the Smith-Waterman algorithm compares segments of all possible lengths and optimizes a similarity measure without looking at any total sequence; cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for replacing a numeric value, wherein a numeric value to use as a replacement depends on any flags or preferences that were set to use a default, the average, a minimum, a maximum, or a null; sampling cleaned, raw-data from the flat-data records in the computer data storage mechanism with an apparatus for executing an algorithm that tests if data are supervised, and if so, that creates a plurality of individual data sets for each class with a stratified selection as needed, and then testing if a selected class is abnormal or uncharacteristic, and if not, down-sampling and producing sampled records of the classes and splitting any remaining data into separate training sets, separate test sets, and separate blind sets all then stored in the computer data storage mechanism for later use in subsequent steps to train a predictive model; if the test for each record of each class in supervised data is abnormal or uncharacteristic, then skipping a down-sampling for that instance; and if in a previous step the cleaned, raw-data from the flat-data records in the computer data storage mechanism was determined by the apparatus for executing an algorithm that tests if data are supervised are, in fact, unsupervised, then down-sampling all records and splitting a remaining a sampled record data into a separate a training set, a separate test set, and a separate blind set for later use in subsequent steps to train a predictive model.
2 . A method that improves the training of predictive models, comprising:
converting and transforming a variety of inconsistent and incoherent supervised and unsupervised training data for predictive models received by a network server as electronic data files, and storing that in a computer data storage mechanism, and then into another single, error-free, uniformly formatted record file stored in the computer data storage mechanism with an apparatus for executing a data integrity analysis algorithm that harmonizes a range of supervised and unsupervised training data into flat-data records in which every field of every record file is modified to be coherent and well-populated with information.
3 . The method of claim 2 , further comprising:
comparing and correcting any data values in each data field in the inconsistent and incoherent supervised and unsupervised training data according to a user-service consumer preference and a predefined data dictionary of valid data values with an apparatus for executing an algorithm that substitutes data values in the data fields of incoming supervised and unsupervised training data with at least one value representing a minimum, a maximum, a null, an average, and a default.
4 . The method of claim 2 , further comprising:
discerning the context of any text included in the inconsistent and incoherent supervised and unsupervised training data with an apparatus for executing a contextual dictionary algorithm that employs a thesaurus of alternative contexts of ambiguous words for find a common context denominator, and to then record the context determined into the computer data storage mechanism for later access by a predictive model.
5 . The method of claim 2 , further comprising:
cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for executing an algorithm for cleaning up raw data in stored data records, field-by-field, record-by-record in which some types of fields are restricted in what is legal or allowed, and includes fetching raw data from the computer data storage mechanism and testing each field if a data value reported is numeric or symbolic, and if numeric, a data dictionary is used to see if such data value is previously listed as valid, and if symbolic, using another data dictionary to see if such data value is listed there as valid.
6 . The method of claim 2 , further comprising:
cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for executing a Smith-Waterman algorithm for a local-sequence alignment and to determine if there are any similar regions between two strings or sequences, and in which a consistent, coherent terminology is then enforceable in each data field without data loss, and in which the Smith-Waterman algorithm compares segments of all possible lengths and optimizes a similarity measure without looking at any total sequence.
7 . The method of claim 2 , further comprising:
cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for replacing a numeric value, wherein a numeric value to use as a replacement depends on any flags or preferences that were set to use a default, the average, a minimum, a maximum, or a null.
8 . The method of claim 2 , further comprising:
sampling cleaned, raw-data from the flat-data records in the computer data storage mechanism with an apparatus for executing an algorithm that tests if data are supervised, and if so, that creates a plurality of individual data sets for each class with a stratified selection as needed, and then testing if a selected class is abnormal or uncharacteristic, and if not, down-sampling and producing sampled records of the classes and splitting any remaining data into separate training sets, separate test sets, and separate blind sets all then stored in the computer data storage mechanism for later use in subsequent steps to train a predictive model; and if the test for each record of each class in supervised data is abnormal or uncharacteristic, then skipping a down-sampling for that instance.
9 . The method of claim 8 , further comprising:
if in a previous step the cleaned, raw-data from the flat-data records in the computer data storage mechanism was determined by the apparatus for executing an algorithm that tests if data are supervised are, in fact, unsupervised, then down-sampling all records and splitting a remaining a sampled record data into a separate a training set, a separate test set, and a separate blind set for later use in subsequent steps to train a predictive model.Join the waitlist — get patent alerts
Track US2016078367A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.