US2006149821A1PendingUtilityA1

Detecting spam email using multiple spam classifiers

Assignee: IBMPriority: Jan 4, 2005Filed: Jan 4, 2005Published: Jul 6, 2006
Est. expiryJan 4, 2025(expired)· nominal 20-yr term from priority
H04L 51/212G06Q 10/107
46
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method for detecting undesirable emails is disclosed. The method combines input from two or more spam classifiers to provide improved classification effectiveness and robustness. The method's effectiveness is improved over that of any one constituent classifier in the sense that the detection rate is increased and/or the false positive rate is decreased. The method's robustness is improved in the sense that, if spammers temporarily elude any one constituent classifier, the other constituent classifiers will still be likely to catch the spam. The method includes obtaining a score from each of a plurality of constituent spam classifiers by applying them to a given input email. The method further includes obtaining a combined spam score from a combined spam classifier that takes as input the plurality of constituent spam classifier scores, the combined spam classifier being computed automatically in accordance with a specified false-positive vs. false-negative tradeoff. The method further includes identifying the given input email as an undesirable email if the combined spam score indicates that the input e-mail is undesirable.

Claims

exact text as granted — not AI-modified
1 . A method of detecting whether a first e-mail is undesirable, the method comprising: 
 inputting the first e-mail to each of a plurality of constituent spam classifiers;    obtaining at least one score from each of the plurality of constituent spam classifiers indicating the degree to which the first e-mail is deemed spam;    obtaining a combined spam score from a combined spam classifier that takes as input the at least one score from the plurality of constituent spam classifiers, the combined spam classifier being computed automatically in accordance with a false-positive vs. false-negative tradeoff; and    identifying the first e-mail as an undesirable e-mail if the combined spam score indicates that the first e-mail is undesirable.    
   
   
       2 . The method of  claim 1 , wherein the step of computing the combined spam classifier comprises: 
 compiling a labeled e-mail corpus comprising a plurality of e-mails that have been labeled according to a degree to which the plurality of e-mails are deemed to be spam;    computing scores of the plurality of constituent spam classifiers on each e-mail in the labeled e-mail corpus; and    analyzing the computed scores of the plurality of constituent spam classifiers on each e- mail in the labeled e-mail corpus to compute a combined spam classifier that best achieves the specified false-positive vs. false-negative tradeoff.    
   
   
       3 . The method of  claim 1 , wherein the step of computing the combined spam classifier comprises: 
 compiling a labeled e-mail corpus consisting of a plurality of e-mails that have been labeled according to the degree to which the plurality of e-mails are deemed to be spam;    computing scores of the plurality of constituent spam classifiers on each e-mail in the labeled e-mail corpus;    establishing a set of one or more sample false-positive vs. false-negative tradeoffs;    analyzing, for each sample false-positive vs. false-negative tradeoff, the computed scores of the plurality of constituent spam classifiers on each e-mail in the labeled e-mail corpus to compute a set of combined spam classifiers, each of which best achieves a corresponding sample false-positive vs. false-negative tradeoff;    selecting a false-positive vs. false-negative tradeoff; and    computing from the false-positive vs. false-negative tradeoff, a set of sample false- positive vs. false-negative tradeoffs and a set of corresponding best combined classifiers a best combined classifier for the false-positive vs. false-negative tradeoff.    
   
   
       4 . The method of  claim 3 , wherein the false-positive vs. false-negative tradeoffs are specified by penalty functions, and the combined spam classifier associated with a given penalty function is computed by an optimization procedure that yields the combined spam classifier for which the value of the given penalty function is minimal on the labeled e-mail corpus.  
   
   
       5 . The method of  claim 4 , wherein the space of possible classifiers is represented by a set of parameterized weights and basis functions, and the optimization procedure searches the parameterized weight space to identify the combined spam classifier for which the given penalty function is minimal on the labeled e-mail corpus.  
   
   
       6 . The method of  claim 5 , wherein the optimization algorithm is a nonlinear derivative-free optimization algorithm.  
   
   
       7 . The method of  claim 5 , wherein the basis functions are individual output scores of the constituent spam classifiers.  
   
   
       8 . The method of  claim 5 , wherein the basis functions are fixed transformations of individual output scores of the constituent spam classifiers.  
   
   
       9 . The method of  claim 5 , wherein the basis functions are parameterized transformations of individual output scores of the constituent spam classifiers, and parameters are included in the search conducted by the optimization algorithm.  
   
   
       10 . The method of  claim 1 , wherein the combined spam score is a numerical value and the combined spam score is considered to be undesirable if it exceeds a specified threshold.  
   
   
       11 . The method of  claim 1 , wherein the at least one score from each of the plurality of constituent spam classifiers is any one of numerical and categorical and includes an output indicating that a constituent spam classifier is unable to assign a definite score.  
   
   
       12 . The method of  claim 1 , wherein the combined spam classifier is recomputed any one of periodically at a specified time interval, in response to a command, and in response to an automatically generated signal.  
   
   
       13 . The method of  claim 12 , wherein the labeled e-mail corpus is updated to include new labeled e-mail and to delete old labeled e-mail when the combined spam classifier is recomputed.  
   
   
       14 . The method of  claim 12 , wherein the automatically generated signal indicates that one or more of the plurality of constituent spam classifiers has changed significantly due to adaptation.  
   
   
       15 . The method of  claim 12 , wherein the automatically generated signal indicates that one or more of the plurality of constituent spam classifiers is performing poorly.  
   
   
       16 . The method of  claim 3 , wherein the false-positive vs. false-negative tradeoff is determined by displaying to a user a set of pairs of estimated false-positive and false-negative rates and allowing the user to select one of the pairs.  
   
   
       17 . The method of  claim 4 , wherein the penalty functions are parameterized by a single parameter that establishes a ratio between a penalty for false positives and a penalty for false negatives.  
   
   
       18 . A method of detecting whether a first e-mail is undesirable, the method comprising the steps of: 
 inputting the first e-mail to a classifier;    obtaining from the classifier a classification of the first e-mail, wherein a range of classifications includes a first classification indicating that the first e-mail cannot be classified as either spam or non-spam; and    taking an action if the first e-mail is classified under the first classification.    
   
   
       19 . The method of  claim 18 , wherein the action comprises: 
 inputting the first e-mail to a second classifier;    obtaining from the second classifier a classification of the first e-mail; and    taking an action if the first e-mail is classified under the first classification.    
   
   
       20 . The method of  claim 18 , wherein the action comprises: 
 placing the first e-mail in a waiting queue; and    re-evaluating the first e-mail at a later time, the re-evaluating comprising: 
 inputting the first e-mail to a second classifier;  
 obtaining from the second classifier a classification of the first e-mail; and  
 taking an action if the first e-mail is classified under the first classification.  
   
   
   
       21 . A method for detecting undesirable e-mail, the method comprising: 
 inputting a first e-mail to each of a plurality of constituent spam classifiers;    obtaining at least one score from each of the plurality of constituent spam classifiers indicating the degree to which the first e-mail is deemed spam;    obtaining a combined spam score from a combined spam classifier that takes as input the at least one score from each of the plurality of constituent spam classifiers, at least one of the plurality of constituent spam classifiers being a member of a similarity-detection family; and    identifying the first e-mail as an undesirable e-mail if the combined spam score indicates that the first e-mail is undesirable.    
   
   
       22 . A computer readable medium including computer instructions for detecting whether a first e-mail is undesirable, the computer instructions including instructions for: 
 inputting the first e-mail to each of a plurality of constituent spam classifiers;    obtaining at least one score from each of the plurality of constituent spam classifiers indicating the degree to which the first e-mail is deemed spam;    obtaining a combined spam score from a combined spam classifier that takes as input the at least one score from the plurality of constituent spam classifiers, the combined spam classifier being computed automatically in accordance with a false-positive vs. false-negative tradeoff; and    identifying the first e-mail as an undesirable e-mail if the combined spam score indicates that the first e-mail is undesirable.    
   
   
       23 . An information processing system for detecting whether a first e-mail is undesirable, comprising: 
 a processor configured for: 
 inputting the first e-mail to each of a plurality of constituent spam classifiers;  
 obtaining at least one score from each of the plurality of constituent spam classifiers indicating the degree to which the first e-mail is deemed spam;  
 obtaining a combined spam score from a combined spam classifier that takes as input the at least one score from the plurality of constituent spam classifiers, the combined spam classifier being computed automatically in accordance with a false-positive vs. false-negative tradeoff; and  
 identifying the first e-mail as an undesirable e-mail if the combined spam score indicates that the first e-mail is undesirable.

Join the waitlist — get patent alerts

Track US2006149821A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.