Systems and methods for providing a content item database and identifying content items
Abstract
Systems and methods are provided for identifying unsolicited or unwanted electronic communications, such as spam. The disclosed embodiments also encompass systems and methods for selecting content items from a content item database. Consistent with certain embodiments, computer-implemented systems and methods may use a clustering based statistical content matching anti-spam algorithm to identify and filter spam. Such a anti-spam algorithm may be implemented to determine a degree of similarity between an incoming e-mail with a collection of one or more spam e-mails stored in a database. If the degree of similarity exceeds a predetermined threshold, the incoming e-mail may be classified as spam. Further, in accordance with other embodiments, systems and methods may be provided to determine a degree of similarity between a query or search string from a user and content items stored in a database. If the degree of similarity exceeds a predetermined threshold, the content item from the database may be identified as a content item that matches the query or search string provided by the user.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method for clustering a content item database, the method comprising:
assigning a content item in the content item database to a first cluster from among a plurality of clusters; identifying representative content items for each of the plurality of clusters; computing a mean vector for each of the plurality of clusters; computing a first distance between the content item and the mean vector for each of the plurality of clusters; reassigning the content item from the first cluster to a second cluster, if the mean vector for the second cluster has the smallest first distance from the content item; computing a second distance between the content item and a representative content item for a third cluster; and reassigning the content item from the second cluster to the third cluster, if the representative content item for the third cluster has the smallest distance from the content item.
2 . The computer-implemented method of claim 1 , wherein the first distance is a cosine distance, and the content item is reassigned from the first cluster to the second cluster, if the mean vector of the second cluster has the largest cosine distance from the content item.
3 . The computer-implemented method of claim 1 , wherein the second distance is a cosine distance and the content item is reassigned from the second cluster to the third cluster if the representative content item of the third cluster has the largest cosine distance from the content item.
4 . The computer-implemented method of claim 1 , wherein the content item comprises at least one of an e-mail, an instant message, a chat message, a text message, a SMS message, a paging communication, a blog post, and a news item.
5 . A computer-implemented method for clustering a content item database, the method comprising:
assigning a content item in the content item database to a first cluster from among a plurality of clusters; identifying representative content items for each of the plurality of clusters; computing a mean vector for each of the plurality of clusters; computing a first degree of similarity between the content item and the mean vector of a second cluster from among the plurality of clusters; reassigning the content item to the second cluster, if the first degree of similarity exceeds a first threshold; computing a second degree of similarity between the content item and a representative content item for a third cluster; and reassigning the content item from the second cluster to the third cluster if the second degree of similarity exceeds a second threshold.
6 . The computer-implemented method of claim 5 , wherein determining the first degree of similarity comprises computing a first easy signature by performing the steps of:
processing the content item by changing an upper-case letter into a lower-case letter and removing a space; creating a first set of tokens from the processed content item, each token having a predetermined length and overlapping a previous token by including one or more characters from the previous token; calculating a first total as a number of tokens in the first set of tokens; determining a number of common tokens present in the first set of tokens and a second set of tokens corresponding to the mean vector of the second cluster; calculating a second total as a number of tokens in the second set of tokens; and determining the first easy signature as a ratio of the number of common tokens and a sum of the first total and the second total.
7 . The computer-implemented method of claim 5 , wherein determining the second degree of similarity comprises computing a second easy signature by performing the steps of:
determining a number of common tokens present in the first set of tokens and the third set of tokens corresponding to the representative content item of the third cluster; calculating a third total as a number of tokens in the third set of tokens; and determining the second easy signature as a ratio of the number of common tokens and a sum of the first total and the third total.
8 .- 24 . (canceled)
25 . A system comprising:
at least one processor; and a memory storing instruction that, when executed by the at least one process, cause the processor to perform operations comprising:
assigning a content item in the content item database to a first cluster from among a plurality of clusters;
identifying representative content items for each of the plurality of clusters;
computing a mean vector for each of the plurality of clusters;
computing a first distance between the content item and the mean vector for each of the plurality of clusters;
reassigning the content item from the first cluster to a second cluster, if the mean vector for the second cluster has the smallest first distance from the content item;
computing a second distance between the content item and a representative content item for a third cluster; and
reassigning the content item from the second cluster to the third cluster, if the representative content item for the third cluster has the smallest distance from the content item.
26 . The system of claim 25 , wherein the first distance is a cosine distance, and the content item is reassigned from the first cluster to the second cluster, if the mean vector of the second cluster has the largest cosine distance from the content item.
27 . The system of claim 25 , wherein the second distance is a cosine distance and the content item is reassigned from the second cluster to the third cluster if the representative content item of the third cluster has the largest cosine distance from the content item.
28 . The system of claim 25 , wherein the content item comprises at least one of an e-mail, an instant message, a chat message, a text message, a SMS message, a paging communication, a blog post, and a news item.
29 . A system comprising:
at least one processor; and a memory storing instruction that, when executed by the at least one process, cause the processor to perform operations comprising:
assigning a content item in the content item database to a first cluster from among a plurality of clusters;
identifying representative content items for each of the plurality of clusters;
computing a mean vector for each of the plurality of clusters;
computing a first degree of similarity between the content item and the mean vector of a second cluster from among the plurality of clusters;
reassigning the content item to the second cluster, if the first degree of similarity exceeds a first threshold;
computing a second degree of similarity between the content item and a representative content item for a third cluster; and
reassigning the content item from the second cluster to the third cluster if the second degree of similarity exceeds a second threshold.
30 . The system of claim 29 , wherein determining the first degree of similarity comprises computing a first easy signature by performing the steps of:
processing the content item by changing an upper-case letter into a lower-case letter and removing a space; creating a first set of tokens from the processed content item, each token having a predetermined length and overlapping a previous token by including one or more characters from the previous token; calculating a first total as a number of tokens in the first set of tokens; determining a number of common tokens present in the first set of tokens and a second set of tokens corresponding to the mean vector of the second cluster; calculating a second total as a number of tokens in the second set of tokens; and determining the first easy signature as a ratio of the number of common tokens and a sum of the first total and the second total.
31 . The system of claim 29 , wherein determining the second degree of similarity comprises computing a second easy signature by performing the steps of:
determining a number of common tokens present in the first set of tokens and the third set of tokens corresponding to the representative content item of the third cluster; calculating a third total as a number of tokens in the third set of tokens; and determining the second easy signature as a ratio of the number of common tokens and a sum of the first total and the third total.Join the waitlist — get patent alerts
Track US2015142809A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.