Method of operating artificial intelligence machines to improve predictive model training and performance
Abstract
A method of improving the training and performance of predictive models. A first method of operating an artificial intelligence machine produces predictive model language documents describing improved predictive models that generate better business decisions from raw data record inputs. A second method of operating an artificial intelligence machine including processors for predictive model algorithms produces and outputs better business decisions from raw data record inputs. Both methods enrich the raw data records their processors are fed by deleting data fields with data values that have little benefit in decision making, and that derive and add new data fields from information sources then available that do benefit in the decision making of the artificial intelligence machine through improved accuracies of prediction.
Claims
exact text as granted — not AI-modified1 . A method of operating an artificial intelligence machine to improve their decisions from included predictive models, comprising:
deleting with at least one processor a selected data field and any data values contained in the selected data field from each of a first series of data training records stored in a memory of the artificial intelligence machine to exclude each data field in the first series of data training records that has more than a threshold number of random data values, or that has only one repeating data value, or that has too small a Shannon entropy, and using any information gained to select the most useful data fields, and then transforming a surviving number of data fields in all the first series of data training records into a corresponding reduced-field series of data training records stored in the memory of the artificial intelligence machine; adding with the at least one processor a new derivative data field to all the reduced-field series of data training records stored in the memory and initializing each added new derivative data field with a new data value, and including an apparatus for executing an algorithm to either change real scaler numeric data values into fuzzy values, or if symbolic, to change a behavior group data value, and testing that a minimum number of data fields survive, and if not, then to generate a new derivative data field and fix within each an aggregation type, a time range, a filter, a set of aggregation constraints, a set of data fields to aggregate, and a recursive level, and then assessing the quality of a newly derived data field by testing it with a test set of data, and then transforming the results into an enriched-field series of data training records stored in the memory of the artificial intelligence machine; verifying with the at least one processor that each predictive model if trained with the enriched-field series of data training records stored in the memory produces decisions having fewer errors than the same predictive model trained only with the first series of data training records; recording a data-enrichment descriptor into the memory to include an identity of selected data fields in a data training record format of the first series of data training records that were subsequently deleted, and which newly derived data fields were subsequently added, and how each newly derived data field was derived and from which information sources; causing the at least one processor of the artificial intelligence machine to start extracting decisions from a new series of data records of new events by receiving and storing the new series of data records in the memory of the artificial intelligence machine; causing the at least one processor to fetch the data-enrichment descriptor and use it to select which data fields to delete and then deleting all the data values included in the selected data fields from each of a new series of data records of new events; wherein, each data field deleted matches a data field in the first series of data training records had more than a threshold number of random data values, or that had only one repeating data value, or that had too small a Shannon entropy; adding with the at least one processor a new derivative data field to each record of the new series of data records stored in the memory according to the data-enrichment descriptor, and initializing each added new derivative data field with a new data value stored in the memory; wherein, each new derivative data field added matches a new derivative data field added to the enriched-field series of data training records in which real scaler numeric data values were changed into fuzzy values, or if symbolic, were changed into a behavior group data value stored in the memory, and were tested that a minimum number of data fields survive, and if not, then that generated a new derivative data field and fixed within each an aggregation type, a time range, a filter, a set of aggregation constraints, a set of data fields to aggregate, and a recursive level; and producing and outputting a series of predictive decisions with the at least one processor that operates at least one predictive model algorithm derived from one originally built and trained with records having a same record format described by the data-enrichment descriptor and stored in the memory of the artificial intelligence machine.
2 . A method of operating an artificial intelligence machine to produce predictive model language documents describing improved predictive models that generate better business decisions from raw data record inputs, comprising:
deleting with at least one processor a selected data field and any data values contained in the selected data field from each of a first series of data records stored in a memory of the artificial intelligence machine to exclude each data field in the first series of data records that has more than a threshold number of random data values, or that has only one repeating data value, or has too small a Shannon entropy, and then transforming a surviving number of data fields in all the first series of data records into a corresponding reduced-field series of data records stored in the memory of the artificial intelligence machine; adding with the at least one processor a new derivative data field to all the reduced-field series of data records stored in the memory of the artificial intelligence machine and initializing each added new derivative data field with a new data value, and including an apparatus for executing an algorithm to either change real scaler numeric data values into fuzzy values, or if symbolic, to change a behavior group data value, and testing that a minimum number of data fields survive, and if not, then to generate a new derivative data field and fix within each an aggregation type, a time range, a filter, a set of aggregation constraints, a set of data fields to aggregate, and a recursive level, and then assessing the quality of a newly derived data field by testing it with a test set of data, and then transforming the results into an enriched-field series of data records stored in the memory of the artificial intelligence machine; and verifying with the at least one processor that a predictive model trained with the enriched-field series of data records stored in the memory of the artificial intelligence machine produces more accurate predictions from the artificial intelligence machine having fewer errors than the same predictive model trained only with the first series of data records.
3 . The method of claim 2 , further comprising:
verifying with the at least one processor that a predictive model supplied with a non-training set of the enriched-field series of data records stored in the memory of the artificial intelligence machine produces more accurate predictions with fewer errors than the same predictive model fed with data records with unmodified data fields.
4 . The method of claim 2 , further comprising:
recording as a data-enrichment descriptor into the memory of the artificial intelligence machine including the at least one processor an identity of any data fields in a data record format of the first series of data records that were subsequently deleted, and which newly derived data fields were subsequently added, and how each newly derived data field was derived and from which information sources; and passing along the data-enrichment descriptor with the at least one processor information stored in the memory of the artificial intelligence machine to an artificial intelligence machine including processors for predictive model algorithms to produce and output better business decisions from its own feed of new events as raw data record inputs stored in the memory of the artificial intelligence machine.
5 . A method of operating an artificial intelligence machine including processors for predictive model algorithms that produces and that outputs better business decisions from a new series of data records of new events as raw data record inputs, comprising:
recovering with at least one processor a recording of a data-enrichment descriptor stored in a memory of the artificial intelligence machine including an identity of any data fields in a data record format of a series of data records that were subsequently deleted by an artificial intelligence machine including processors for predictive model building, and which of any newly derived data fields were subsequently added, and how each newly derived data field was derived and from which information sources; accepting a new series of data records of new events with the artificial intelligence machine including at least one processor to receive and store records in the memory of the artificial intelligence machine; deleting with the at least one processor all data fields and all data values contained in the data fields from each of a new series of data records of new events, stored in the memory of the artificial intelligence machine, according to the data-enrichment descriptor; adding with the at least one processor a new derivative data field to each record of the new series of data records stored in the memory of the artificial intelligence machine according to the data-enrichment descriptor, and initializing each added new derivative data field with a new data value stored in the memory of the artificial intelligence machine; and producing and outputting a series of predictive decisions with the at least one processor that operates at least one predictive model algorithm derived from one originally built and trained with records having a same record format described by the data-enrichment descriptor and stored in the memory of the artificial intelligence machine.
6 . The method of claim 5 which includes causing the at least one processor in the step of deleting to:
exclude each data field stored in the memory of the artificial intelligence machine that has more than a threshold number of random data values, or that has only one repeating data value, or that has too small a Shannon entropy, and then transforming a surviving number of data fields into a corresponding reduced-field series of data records stored in the memory of the artificial intelligence machine.
7 . The method of claim 6 which includes causing the at least one processor in the step of adding to:
add a new derivative data field to a reduced-field series of data records stored in the memory of the artificial intelligence machine and initialize each added new derivative data field with a new data value, and to either change real scaler numeric data values into fuzzy values, or if symbolic, to change a behavior group data value stored in the memory of the artificial intelligence machine, and testing that a minimum number of data fields survive in that stored in the memory of the artificial intelligence machine, and if not, then to generate a new derivative data field and fix within each an aggregation type, a time range, a filter, a set of aggregation constraints, a set of data fields to aggregate, and a recursive level, and which the quality of each newly derived data field was test, and then transforming the results into an enriched-field series of data records stored in the memory of the artificial intelligence machine.Join the waitlist — get patent alerts
Track US2016071017A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.