Data Mining Technique With Maintenance of Ancestry Counts
Abstract
Roughly described, a computer-implemented evolutionary data mining system includes a memory storing a candidate gene database in which each candidate individual has a respective fitness estimate; a gene pool processor which tests individuals from the candidate gene pool on training data and updates the fitness estimate associated with the individuals in dependence upon the tests; and a gene harvesting module for deploying selected individuals from the gene pool, wherein the gene pool processor includes a competition module which selects individuals for discarding in dependence upon their updated fitness estimate. The system maintains the ancestry count for each of the candidate individuals, and may use this information to adjust the competition among the individuals, to adjust the selection of individuals for further procreation, and/or for other purposes.
Claims
exact text as granted — not AI-modified1 . A computer-implemented data mining method, for use with a data mining training database containing training data, comprising the steps of:
providing a system having a candidate gene database identifying a candidate pool of candidate individuals, each candidate individual identifying a plurality of conditions and at least one corresponding proposed output in dependence upon the conditions, each candidate individual further having associated therewith an indication of a respective fitness estimate and an ancestry count, wherein ancestry count corresponds to a number of procreation events that occurred in a procreation history of each candidate individual; testing on the training data by a gene testing module of the system, each individual in a testing subset of at least one of the candidate individuals, each individual in the testing subset undergoing at least one trial, each trial applying the conditions of the respective individual to the training data to propose a result; calculating by the gene testing module, an estimated fitness of each of the individuals in the testing subset in dependence upon the training data and the results proposed by the individual in the step of testing; and populating an elitist pool with elite individuals by a competition module in accordance with (i) a predetermined requirement for fitness estimate and (ii) a maximum number of allowed candidate individuals; and further wherein the competition module adjusts the fitness estimate of each individual in dependence upon the individual's ancestry count and selects an individual with a higher fitness estimate over an individual with a lower fitness estimate for inclusion in the elitist pool.
2 . The computer-implemented data mining method of claim 1 , wherein the competition module applies a handicap of a fixed percentage against individuals whose ancestry count exceeds a predetermined number of generations.
3 . The computer-implemented data mining method of claim 1 , wherein in populating the elitist pool with elite individuals, the competition module further (iii) compares an experience level of each individual to a predetermined threshold experience level, wherein experience level is determined by a total number of trials the individual has undergone.
4 . The computer-implemented data mining method of claim 2 , wherein the handicap applied to each given one of the individuals by the competition module varies non-decreasingly as a function of the ancestry count of the given individual.
5 . A computer-implemented data mining method for use with a data mining training database containing a plurality of data samples, comprising:
providing a system having a candidate gene database identifying a candidate pool of candidate individuals, each candidate individual identifying a plurality of conditions and at least one corresponding proposed output in dependence upon the conditions, each candidate individual further having associated therewith an indication of a respective fitness estimate and an ancestry count, wherein ancestry count corresponds to a number of procreation events that occurred in a procreation history of each candidate individual; performing a procreation step by a procreation module of a gene pool processor of forming new candidate individuals in the candidate pool of candidate individuals, wherein each new individual is related to one or more previously existing parent candidate individuals; testing by a gene testing module of the gene pool processor each individual in a testing subset of at least one of the candidate individuals, wherein the testing subset includes at least one new candidate individual, each of the tests applying the conditions of the respective individual to a respective subset of the data samples in a training database to propose a result, each individual in the testing subset being tested on at least one data sample and at least one of the individuals in the testing subset being tested on more than one data sample; calculating by a competition module of the gene pool processor an overall fitness estimate for each of the individuals in the testing subset, in dependence upon the results proposed by the respective individual when the conditions of the respective individual were applied to the respective subset of the data samples; and selecting one or more individuals for discarding from the candidate pool by the competition module in accordance with (i) a predetermined requirement for fitness estimate and (ii) a maximum number of allowed candidate individuals; further wherein the competition module adjusts the fitness estimate of each individual in dependence upon the individual's ancestry count and discards an individual with a lower fitness estimate over an individual with a higher fitness estimate; and
harvesting by a gene harvesting module selected ones of non-discarded individuals from the candidate pool of candidate individuals for deployment of selected individuals in a predetermined production environment, wherein candidate individuals are selected in dependence upon comparisons among their respective ancestry counts.
6 . The computer-implemented data mining method of claim 5 , the procreation module selects the one or more previously existing parent candidate individuals for forming a new candidate individual in accordance with an ancestry count for each of the one or more previously existing parent candidate individuals.
7 . A computer-implemented data mining method for use with a client-server architecture, comprising:
providing a server having a candidate gene database identifying a candidate pool of candidate individuals, each candidate individual identifying a plurality of conditions and at least one corresponding proposed output in dependence upon the conditions, each candidate individual further having associated therewith an indication of a respective fitness estimate and an ancestry count, wherein ancestry count corresponds to a number of procreation events that occurred in a procreation history of each candidate individual; delegating by the server testing subsets containing at least one of the candidate individuals to individual clients, wherein the testing by the individual clients includes:
testing each individual in a delegated testing subset on training data by a gene testing module of the client, each individual in the delegated testing subset undergoing at least one trial, each trial applying the conditions of the respective individual to the training data to propose a result;
calculating by the gene testing module, an estimated fitness of each of the individuals in the delegated testing subset in dependence upon the training data and the results proposed by the individual in the step of testing; and
providing tested individuals, including evaluation data, to the server, wherein the evaluation data includes the calculated estimated fitness and ancestry count;
adjusting the calculated estimated fitness for received tested individuals by a competition module on the server in dependence upon the individual's ancestry count; and updating the candidate pool by the competition module to include individuals in accordance with (i) a predetermined requirement for fitness estimate and (ii) a maximum number of allowed candidate individuals, wherein the competition module selects an individual with a higher fitness estimate over an individual with a lower fitness estimate for inclusion in the candidate pool.
8 . The computer-implemented data mining method of claim 7 , further comprising:
performing a procreation step by a procreation module of the gene pool processor of each client of forming new candidate individuals in the candidate pool of candidate individuals, wherein each new individual is related to one or more previously existing parent candidate individuals; testing by a gene testing module of the gene pool processor of each client each individual in a testing subset of at least one of the candidate individuals, wherein the testing subset includes at least one new candidate individual, each of the tests applying the conditions of the respective individual to a respective subset of the data samples in a training database to propose a result, each individual in the testing subset being tested on at least one data sample and at least one of the individuals in the testing subset being tested on more than one data sample; providing tested procreated individuals, including evaluation data, to the server, wherein the evaluation data includes the calculated estimated fitness and ancestry count; calculating by a competition module on the server an overall fitness estimate for each of the procreated individuals in the testing subset, in dependence upon the results proposed by the respective procreated individual when the conditions of the respective procreated individual were applied to the respective subset of the data samples; and selecting one or more procreated individuals for discarding from the candidate pool by the competition module in accordance with (i) a predetermined requirement for fitness estimate and (ii) a maximum number of allowed candidate individuals; further wherein the competition module adjusts the fitness estimate of each procreated individual in dependence upon the individual's ancestry count and discards an individual with a lower fitness estimate over an individual with a higher fitness estimate; and harvesting by a gene harvesting module selected ones of non-discarded procreated individuals from the candidate pool of candidate individuals for deployment of selected procreated individuals in a predetermined production environment, wherein candidate procreated individuals are selected in dependence upon comparisons among their respective ancestry counts.Join the waitlist — get patent alerts
Track US2019220751A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.