US2010293116A1PendingUtilityA1

Url and anchor text analysis for focused crawling

Assignee: FENG SHI CONGPriority: Nov 8, 2007Filed: Nov 8, 2007Published: Nov 18, 2010
Est. expiryNov 8, 2027(~1.3 yrs left)· nominal 20-yr term from priority
G06F 16/951G06F 16/9566
40
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods of URL and anchor text analysis for focused crawling are disclosed. In an exemplary embodiment, a method may include training a focused crawler by: obtaining a training set of at least URL's or anchor text for a website, computing a score for the training set, and extracting a plurality of features of the training set, and computing a score for each of the plurality of features. The features identify key information contained in the website. The method may also include executing a trained focused crawler on other websites.

Claims

exact text as granted — not AI-modified
1 . A method of Uniform Resource Locator (URL) and anchor text analysis for focused crawling, comprising:
 training a focused crawler by:
 obtaining a training set for a website; 
 computing a score for the training set of at least URL's or anchor text; 
 extracting a plurality of features of the training set, the features identifying key information contained in the website; and 
 computing a score for each of the plurality of features; and 
   executing a trained focused crawler on other websites.   
     
     
         2 . The method of  claim 1  wherein obtaining the training set is by downloading a plurality of complete websites related to a type of website for focused crawling. 
     
     
         3 . The method of  claim 1  wherein a higher score indicates the URL refers to a target page, or the URL leads quickly to a target page. 
     
     
         4 . The method of  claim 1  wherein computing the score is by manual labeling, or by automatic labeling using a software classifier based on content of each web page in the website, or by link structure analysis. 
     
     
         5 . The method of  claim 1  wherein features include phrases, multiple words concatenated into one phrase and separated into individual features, stemmed words, position of a phrase, a co-appearance relationship, relative positions, or patterns. 
     
     
         6 . The method of  claim 1  wherein the score of a feature satisfies the following criteria: each occurrence of a feature with a positive score makes a positive contribution to the score of the feature, and each occurrence of a feature with a negative score makes a negative contribution to the score of the feature, and neutral features have a neutral score. 
     
     
         7 . The method of  claim 1  wherein more common features result in higher scores and more dispersed features result in lower scores. 
     
     
         8 . The method of  claim 1  wherein executing a trained focused crawler on other websites is by:
 extracting features from each other website; and   determining whether to download a web page based on the score.   
     
     
         9 . The method of  claim 8  wherein the determination is made using a threshold. 
     
     
         10 . The method of  claim 9  wherein the threshold is after a predetermined number of pages are downloaded. 
     
     
         11 . The method of  claim 9  wherein the threshold is after a predetermined time has passed. 
     
     
         12 . A system comprising:
 a training module operating to obtain a training set for a website, compute a score for the training set, and extract a plurality of features of the training set, the features identifying key information contained in the website; and   an execution module operating to compute a score for each of the plurality of features, and crawl other websites.   
     
     
         13 . The system of  claim 12  wherein features include phrases, multiple words concatenated into one phrase and separated into individual features, stemmed words, position of a phrase, a co-appearance relationship, relative positions, or patterns. 
     
     
         14 . The system of  claim 12  wherein the score of a feature satisfies the following criteria: each occurrence of a feature with a positive score makes a positive contribution to the score of the feature, and each occurrence of a feature with a negative score makes a negative contribution to the score of the feature, and neutral features have a neutral score. 
     
     
         15 . The system of  claim 12  wherein more common features result in higher scores. 
     
     
         16 . The system of  claim 12  wherein more dispersed features result in lower scores. 
     
     
         17 . The system of  claim 12  wherein executing a trained focused crawler on other websites is by:
 extracting features from each other website; and   determining whether to download a web page based on the score.   
     
     
         18 . The system of  claim 17  wherein the determination is made using a threshold. 
     
     
         19 . The system of  claim 18  wherein the threshold is after a predetermined number of pages are downloaded or after a predetermined time has passed. 
     
     
         20 . A system for focused crawling using Uniform Resource Locator (URL) and anchor text analysis, comprising:
 means for training a focused crawler by obtaining a training set of at least URLs or anchor text for a website, computing a score for the training set, and extracting a plurality of features of the training set, and computing a score for each of the plurality of features, wherein the features identify key information contained in the website; and   means for executing a trained focused crawler on other websites.

Join the waitlist — get patent alerts

Track US2010293116A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.