US2003023571A1PendingUtilityA1

Enhancing knowledge discovery using support vector machines in a distributed network environment

Priority: Mar 7, 1997Filed: Aug 21, 2002Published: Jan 30, 2003
Est. expiryMar 7, 2017(expired)· nominal 20-yr term from priority
G06F 18/214G06N 20/00G06F 18/2411G06N 20/10
45
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method for enhancing knowledge discovery from data using a learning machine in general and a support vector machine in particular in a distributed network environment. A customer may transmit training data, test data and live data to a vendor's server from a remote source, via a distributed network. The customer may also transmit to the server identification information such as a user name, a password and a financial account identifier. The training data, test data and live data may be stored in a storage device. Training data may then be pre-processed in order to add meaning thereto. Pre-processing data may involve transforming the data points and/or expanding the data points. By adding meaning to the data, the learning machine is provided with a greater amount of information for processing. With regard to support vector machines in particular, the greater the amount of information that is processed, the better generalizations about the data that may be derived. The learning machine is therefore trained with the pre-processed training data and is tested with test data that is pre-processed in the same manner. The test output from the learning machine is post-processed in order to determine if the knowledge discovered from the test data is desirable. Post-processing involves interpreting the test output into a format that may be compared with the test data. Live data is pre-processed and input into the trained and tested learning machine. The live output from the learning machine may then be post-processed into a computationally derived alphanumerical classifier for interpretation by a human or computer automated process. Prior to transmitting the alpha numerical classifier to the customer via the distributed network, the server is operable to communicate with a financial institution for the purpose of receiving funds from a financial account of the customer identified by the financial account identifier.

Claims

exact text as granted — not AI-modified
What is claimed is:  
     
         1 . A system for enhancing knowledge discovery using a support vector machine comprising: 
 a server in communication with a distributed network for receiving a training data set, a test data set, a live data set and a financial account identifier from a remote source, the remote source also in communication with the distributed network;    one or more storage devices in communication with the server for storing the training data set and the test data set;    a processor for executing a support vector machine;    the processor further operable for: 
 collecting the training data set from the one or more storage devices,  
 pre-processing the training data set to add meaning to each of a plurality of training data points,  
 inputting the pre-processed training data set into the support vector machine so as to train the support vector machine,  
 in response to training of the support vector machine, collecting the test data set from the database,  
 pre-processing the test data set in the same manner as was the training data set,  
 inputting the test data set into the trained support vector machine in order to test the support vector machine,  
 in response to receiving a test output from the trained support vector machine, collecting the live data set from the one or more storage devices,  
 inputting the live data set into the tested and trained support vector machine in order to process the live data,  
 in response to receiving a live output from the support vector machine, post-processing the live output to derive a computationally based alpha numerical classifier, and  
 transmitting the alphanumerical classifier to the server;  
   wherein the server is further operable for: 
 communicating with a financial institution in order to receive funds from a financial account identified by the financial account identifier, and  
 in response to receiving the funds, transmitting the alphanumerical identifier to the remote source or another remote source.  
   
     
     
         2 . The system of  claim 1 , wherein each training data point comprises a vector having one or more coordinates; and 
 wherein pre-processing the-training data set to add meaning to each training data point comprises: 
 determining that the training data point is dirty; and  
 in response to determining that the training data point is dirty, cleaning the training data point.  
   
     
     
         3 . The system of  claim 2 , wherein cleaning the training data point comprises deleting, repairing or replacing the data point.  
     
     
         4 . The system of  claim 1 , wherein each training data point comprises a vector having one or more original coordinates; and 
 wherein pre-processing the training data set to add meaning to each training data point comprises adding dimensionality to each training data point by adding one or more new coordinates to the vector.    
     
     
         5 . The system of  claim 4 , wherein the one or more new coordinates added to the vector are derived by applying a transformation to one or more of the original coordinates.  
     
     
         6 . The system of  claim 4 , wherein the transformation is based on expert knowledge.  
     
     
         7 . The system of  claim 4 , wherein the transformation is computationally derived.  
     
     
         8 . The system of  claim 4 , wherein the training data set comprises a continuous variable; and 
 wherein the transformation comprises optimally categorizing the continuous variable of the training data set.    
     
     
         9 . The system of  claim 1 , wherein the knowledge to be discovered from the data relates to a regression or density estimation; 
 wherein the support vector machine produces a training output comprising a continuous variable; and    wherein the processor is further operable for post-processing the training output by optimally categorizing the training output to derive cutoff points in the continuous variable.    
     
     
         10 . The system of  claim 1 , wherein the processor is further operable for: 
 in response to comparing each of the test outputs with each other, determining that none of the test outputs is the optimal solution;    adjusting the different kernels of one or more of the plurality of support vector machines; and    in response to adjusting the selection of the different kernels, retraining and retesting each of the plurality of support vector machines.

Join the waitlist — get patent alerts

Track US2003023571A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.