Classifier Tuning Based On Data Similarities
Abstract
A probabilistic classifier is used to classify data items in a data stream. The probabilistic classifier is trained, and an initial classification threshold is set, using unique training and evaluation data sets (i.e., data sets that do not contain duplicate data items). Unique data sets are used for training and in setting the initial classification threshold so as to prevent the classifier from being improperly biased as a result of similarity rates in the training and evaluation data sets that do not reflect similarity rates encountered during operation. During operation, information regarding the actual similarity rates of data items in the data stream is obtained and used to adjust the classification threshold such that misclassification costs are minimized given the actual similarity rates.
Claims
exact text as granted — not AI-modified1 . A method for adjusting a classification threshold of a data item classifier based on received data items, wherein the data item classifier classifies an incoming data item as a member of a particular class when a comparison of a classification output for the incoming data item to the classification threshold indicates the data item belongs to the particular class, the method comprising:
determining the similarity rate for unique data items in the received data items; determining a threshold value for the classification threshold that reduces misclassification costs based on, at least in part, the similarity rate for unique data items in the received data items; and setting the classification threshold to the threshold value.
2 . The method of claim 1 further comprising:
for at least one received data item, obtaining a classification output indicative of whether or not the data item belongs to the particular class; and wherein the threshold value for the classification threshold that reduces misclassification costs is determined also based on, at least in part, the classification output of the at least one data item.
3 . The method of claim 2 wherein obtaining a classification output for the at least one data item comprises:
obtaining feature data for the data item by determining whether the data item has a predefined set of features; inputting the feature data into a probabilistic classifier to obtain a probability measure; and producing a classification output based on the probability measure.
4 . The method of claim 1 further comprising:
receiving a class indication for the at least one data item; and wherein the value for the classification threshold that reduces misclassification costs is determined also based on, at least in part, the class indication of the at least one data item.
5 . The method of claim 4 wherein determining a value for the classification threshold comprises:
determining a value that minimizes: L ( v ) = ∑ x P ( x ❘ v ) ( P ( s ❘ x , v ) [ F ( x ) = l ] + cos t ( x , v ) · P ( l ❘ x , v ) [ F ( x ) = s ] ) where v represents a particular similarity rate, P(x|v) is the probability that the particular data item x occurs given the particular similarity rate v, P(s|x,v) is the probability that the data item x is the particular class given the particular similarity rate v, P(l|x,v) is the probability the particular data item x is not the particular class given the particular similarity rate v, [F(x)=s] is equal to one when an e-mail x is classified as a member of the particular class, zero otherwise, [F(x)=l] is equal to one when an e-mail x is not classified as a member of the particular class, zero otherwise, and cost (x,v) represents an assigned cost of misclassifying data items that are not members of the particular class as members of the particular class.
6 . The method of claim 1 wherein determining a threshold value that reduces misclassification costs comprises determining a threshold value that minimizes misclassification costs.
7 . The method of claim 1 wherein the data items are e-mails and the particular class is spam, such that the data item classifier is used to filter out spam e-mail in a set of received e-mails of unknown classification.
8 . The method of claim 7 wherein the misclassification costs depend on varying costs of misclassifying subcategories of non-spam e-mail as spam e-mail.
9 . A computer-usable medium having a computer program embodied thereon for adjusting a classification threshold of a data item classifier based on received data items, wherein the data item classifier classifies an incoming data item as a member of a particular class when a comparison of a classification output for the incoming data item to the classification threshold indicates the data item belongs to the particular class, the computer program comprising instructions for causing a computer to perform the following operations:
determine the similarity rate for unique data items in the received data items; determine a threshold value for the classification threshold that reduces misclassification costs based on, at least in part, the similarity rate for unique data items in the received data items; and set the classification threshold to the threshold value.
10 . The computer-usable medium of claim 9 wherein the computer program further comprises instructions for causing a computer to perform the following operations:
for at least one received data item, obtain a classification output indicative of whether or not the data item belongs to the particular class; and wherein the threshold value for the classification threshold that reduces misclassification costs is determined also based on, at least in part, the classification output of the at least one data item.
11 . The computer-usable medium of claim 10 wherein, to obtain a classification output for the at least one data item, the computer program comprises instructions for causing a computer to:
obtain feature data for the data item by determining whether the data item has a predefined set of features; input the feature data into a probabilistic classifier to obtain a probability measure; and produce a classification output based on the probability measure.
12 . The computer-usable medium of claim 9 wherein the computer program further comprises instructions for causing a computer to perform the following operations:
receive a class indication for the at least one data item; and wherein the value for the classification threshold that reduces misclassification costs is determined also based on, at least in part, the class indication of the at least one data item.
13 . The computer-usable medium of claim 12 wherein, to determine a value for the classification threshold, the computer program comprises instructions for causing a computer to:
determine a value that minimizes: L ( v ) = ∑ x P ( x ❘ v ) ( P ( s ❘ x , v ) [ F ( x ) = l ] + cos t ( x , v ) · P ( l ❘ x , v ) [ F ( x ) = s ] ) where v represents a particular similarity rate, P(x|v) is the probability that the particular data item x occurs given the particular similarity rate v, P(s|x,v) is the probability that the data item x is the particular class given the particular similarity rate v, P(l|x,v) is the probability the particular data item x is not the particular class given the particular similarity rate v, [F(x)=s] is equal to one when an e-mail x is classified as a member of the particular class, zero otherwise, [F(x)=l] is equal to one when an e-mail x is not classified as a member of the particular class, zero otherwise, and cost (x,v) represents an assigned cost of misclassifying data items that are not members of the particular class as members of the particular class.
14 . The computer-usable medium of claim 9 wherein, to determine a threshold value that reduces misclassification costs, the computer program comprises instructions for causing a computer to determine a threshold value that minimizes misclassification costs.
15 . The computer-usable medium of claim 9 wherein the data items are e-mails and the particular class is spam, such that the data item classifier is used to filter out spam e-mail in a set of received e-mails of unknown classification.
16 . The computer-usable medium of claim 15 wherein the misclassification costs depend on varying costs of misclassifying subcategories of non-spam e-mail as spam e-mail.Join the waitlist — get patent alerts
Track US2006190481A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.