US2011213763A1PendingUtilityA1

Web content mining of pair-based data

Assignee: MICROSOFT CORPPriority: Nov 19, 2007Filed: May 4, 2011Published: Sep 1, 2011
Est. expiryNov 19, 2027(~1.3 yrs left)· nominal 20-yr term from priority
G06F 16/951
48
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Described herein is technology for, among other things, mining pair-based data on the web. The technology involves an online pair-based data mining system as well as an offline SVM training system. By subjecting a pair-based input data to the systems, one may grow a pool of pair-based data which share characteristics of the pair-based input data in more efficient manner.

Claims

exact text as granted — not AI-modified
1 . A system comprising:
 a computer;   a parser implemented at least in part by the computer, configured for receiving results of a search based on a seed term, and further configured for generating a snippet set comprising at least one term from the received results that is associated with the seed term based on rules that include the seed term and the at least one term sharing a common number of characters, words, or syllables, being harmonious, having matching parts of speech, and sharing a common writing style; and   a filter implemented at least in part by the computer and configured for generating at least one couplet comprising the seed term and a second term from the snippet set that complies with the rules.   
     
     
         2 . The system of  claim 1  wherein the filter comprises a support vector machine classifier configured for generating, at least in part, the at least one couplet. 
     
     
         3 . The system of  claim 2  wherein the generating comprises keeping positive candidate data and dropping negative candidate data as determined by the support vector machine classifier. 
     
     
         4 . The system of  claim 1  wherein the search is performed via a search engine. 
     
     
         5 . The system of  claim 1  wherein the seed term is in a first language and the at least one term is in a second language. 
     
     
         6 . The system of  claim 1  wherein the filter comprises:
 an identity filter configured for discarding the snippet set in response to the snippet set not including at least the seed term, or 
 a neighbor filter configured for generating term pairs based on the at least one term of the snippet set, or 
 a length filter configured for discarding ones of the term pairs that do not have a same length, or 
 a frequency filter configured for discarding ones of the term pairs having a frequency that is less than a threshold. 
 
     
     
         7 . The system of  claim 1  wherein the at least one couplet is a Chinese couplet. 
     
     
         8 . A method comprising:
 generating, by a computer, a snippet set comprising at least one term from results of a search based on a seed term, the at least one term associated with the seed term based on rules that include the seed term and the at least one term sharing a common number of characters, words, or syllables, being harmonious, having matching parts of speech, and sharing a common writing style; and   generating at least one couplet comprising the seed term and a second term from the snippet set that complies with the rules.   
     
     
         9 . The method of  claim 8  wherein the generating the at least one couplet is performed by a support vector machine classifier. 
     
     
         10 . The method of  claim 9  wherein the generating the at least one couplet comprises keeping positive candidate data and dropping negative candidate data as determined by the support vector machine classifier. 
     
     
         11 . The method of  claim 8  wherein the search is performed via a search engine. 
     
     
         12 . The method of  claim 8  wherein the seed term is in a first language and the at least one term is in a second language. 
     
     
         13 . The method of  claim 8  wherein the generating the at least one couplet comprises:
 discarding the snippet set in response to the snippet set not including at least the seed term, or 
 generating term pairs based on the at least one term of the snippet set, or 
 discarding ones of the term pairs that do not have a same length, or 
 discarding ones of the term pairs having a frequency that is less than a threshold. 
 
     
     
         14 . The method of  claim 8  wherein the at least one couplet is a Chinese couplet. 
     
     
         15 . At least one computer readable medium storing computer-executable instructions that, when executed by a computer, cause the computer to perform method comprising:
 generating a snippet set comprising at least one term from results of a search based on a seed term, the at least one term associated with the seed term based on rules that include the seed term and the at least one term sharing a common number of characters, words, or syllables, being harmonious, having matching parts of speech, and sharing a common writing style; and   generating at least one couplet comprising the seed term and a second term from the snippet set that complies with the rules.   
     
     
         16 . The at least one computer readable medium of  claim 15  wherein the generating the at least one couplet is performed by a support vector machine classifier. 
     
     
         17 . The at least one computer readable medium of  claim 16  wherein the generating the at least one couplet comprises keeping positive candidate data and dropping negative candidate data as determined by the support vector machine classifier. 
     
     
         18 . The at least one computer readable medium of  claim 15  wherein the search is performed via a search engine. 
     
     
         19 . The at least one computer readable medium of  claim 15  wherein the seed term is in a first language and the at least one term is in a second language. 
     
     
         20 . The at least one computer readable medium of  claim 15  wherein the generating the at least one couplet comprises:
 discarding the snippet set in response to the snippet set not including at least the seed term, or 
 generating term pairs based on the at least one term of the snippet set, or 
 discarding ones of the term pairs that do not have a same length, or 
 discarding ones of the term pairs having a frequency that is less than a threshold.

Join the waitlist — get patent alerts

Track US2011213763A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.