US2010293116A1PendingUtilityA1
Url and anchor text analysis for focused crawling
Est. expiryNov 8, 2027(~1.3 yrs left)· nominal 20-yr term from priority
G06F 16/951G06F 16/9566
40
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Systems and methods of URL and anchor text analysis for focused crawling are disclosed. In an exemplary embodiment, a method may include training a focused crawler by: obtaining a training set of at least URL's or anchor text for a website, computing a score for the training set, and extracting a plurality of features of the training set, and computing a score for each of the plurality of features. The features identify key information contained in the website. The method may also include executing a trained focused crawler on other websites.
Claims
exact text as granted — not AI-modified1 . A method of Uniform Resource Locator (URL) and anchor text analysis for focused crawling, comprising:
training a focused crawler by:
obtaining a training set for a website;
computing a score for the training set of at least URL's or anchor text;
extracting a plurality of features of the training set, the features identifying key information contained in the website; and
computing a score for each of the plurality of features; and
executing a trained focused crawler on other websites.
2 . The method of claim 1 wherein obtaining the training set is by downloading a plurality of complete websites related to a type of website for focused crawling.
3 . The method of claim 1 wherein a higher score indicates the URL refers to a target page, or the URL leads quickly to a target page.
4 . The method of claim 1 wherein computing the score is by manual labeling, or by automatic labeling using a software classifier based on content of each web page in the website, or by link structure analysis.
5 . The method of claim 1 wherein features include phrases, multiple words concatenated into one phrase and separated into individual features, stemmed words, position of a phrase, a co-appearance relationship, relative positions, or patterns.
6 . The method of claim 1 wherein the score of a feature satisfies the following criteria: each occurrence of a feature with a positive score makes a positive contribution to the score of the feature, and each occurrence of a feature with a negative score makes a negative contribution to the score of the feature, and neutral features have a neutral score.
7 . The method of claim 1 wherein more common features result in higher scores and more dispersed features result in lower scores.
8 . The method of claim 1 wherein executing a trained focused crawler on other websites is by:
extracting features from each other website; and determining whether to download a web page based on the score.
9 . The method of claim 8 wherein the determination is made using a threshold.
10 . The method of claim 9 wherein the threshold is after a predetermined number of pages are downloaded.
11 . The method of claim 9 wherein the threshold is after a predetermined time has passed.
12 . A system comprising:
a training module operating to obtain a training set for a website, compute a score for the training set, and extract a plurality of features of the training set, the features identifying key information contained in the website; and an execution module operating to compute a score for each of the plurality of features, and crawl other websites.
13 . The system of claim 12 wherein features include phrases, multiple words concatenated into one phrase and separated into individual features, stemmed words, position of a phrase, a co-appearance relationship, relative positions, or patterns.
14 . The system of claim 12 wherein the score of a feature satisfies the following criteria: each occurrence of a feature with a positive score makes a positive contribution to the score of the feature, and each occurrence of a feature with a negative score makes a negative contribution to the score of the feature, and neutral features have a neutral score.
15 . The system of claim 12 wherein more common features result in higher scores.
16 . The system of claim 12 wherein more dispersed features result in lower scores.
17 . The system of claim 12 wherein executing a trained focused crawler on other websites is by:
extracting features from each other website; and determining whether to download a web page based on the score.
18 . The system of claim 17 wherein the determination is made using a threshold.
19 . The system of claim 18 wherein the threshold is after a predetermined number of pages are downloaded or after a predetermined time has passed.
20 . A system for focused crawling using Uniform Resource Locator (URL) and anchor text analysis, comprising:
means for training a focused crawler by obtaining a training set of at least URLs or anchor text for a website, computing a score for the training set, and extracting a plurality of features of the training set, and computing a score for each of the plurality of features, wherein the features identify key information contained in the website; and means for executing a trained focused crawler on other websites.Join the waitlist — get patent alerts
Track US2010293116A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.