US2010287152A1PendingUtilityA1

System, method and computer readable medium for web crawling

Assignee: SUBOTI LLCPriority: May 5, 2009Filed: May 5, 2009Published: Nov 11, 2010
Est. expiryMay 5, 2029(~2.8 yrs left)· nominal 20-yr term from priority
G06F 16/951
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In a web crawler, a URL selection module selects URLs for pages to be downloaded. The URL selection module accesses an interaction data store that stores interaction data for web pages, including interaction data that indicates human interactions with the pages. To reduce the effects of link farms, the URL selection module filters the URLs to select only those URLs that have human interaction histories and provides the selected URLs to a download module for web page downloading.

Claims

exact text as granted — not AI-modified
1 . A method for web crawling comprising:
 determining a plurality of Uniform Resource Locators (URL)s;   determining a subset of the plurality of URLs that have associated interaction data;   selecting at least one URL of the subset; and   downloading a web page corresponding to the at least one selected URL.   
     
     
         2 . The method according to  claim 1  wherein determining the subset comprises:
 accessing an interaction data store that stores interaction data that associates a URL with an interaction with a web page corresponding to the respective URL; and   selecting a URL into the subset if a web page corresponding to the URL has interaction data.   
     
     
         3 . The method according to  claim 2  comprising ranking the subset of URLs using the interaction data. 
     
     
         4 . The method according to  claim 3  wherein the interaction data indicates one or more out-click events from the corresponding web page of an associated URL and wherein ranking the subset of URLs comprises ranking the URLs dependent on the one or more out-click events. 
     
     
         5 . The method according to  claim 4  wherein ranking a URL is dependent on a source content element of the web page prior to an out-click event. 
     
     
         6 . The method according to  claim 5  wherein ranking a URL is dependent on an attention analysis ranking of the source content element. 
     
     
         7 . The method according to  claim 3  wherein selecting at least one URL comprises selecting a highest ranked URL. 
     
     
         8 . The method according to  claim 2  comprising selecting a URL into the subset if a web page corresponding to the URL has interaction data that indicates at least one human dependent interaction with a web page associated with the URL. 
     
     
         9 . The method according to  claim 1  wherein downloading a web page comprises determining content interest of one or more content elements of the web page. 
     
     
         10 . The method according to  claim 9  comprising storing content elements that satisfy a threshold content interest requirement. 
     
     
         11 . The method according to  claim 1  wherein said downloading comprises providing the subset of URLs to a download module. 
     
     
         12 . The method according to  claim 1  further comprising storing the interaction data comprising:
 receiving an event stream from an interaction between a user and a web page;   analyzing the event stream; and   storing the analyzed event stream in association with a URL for the respective web page.   
     
     
         13 . A web crawler comprising:
 at least one Uniform Resource Locator (URL) data store that stores a plurality of URLs;   at least one interaction data store that stores interaction data for a plurality of web pages, the interaction data indicating an interaction between a human and a web page corresponding to a URL;   at least one download module that downloads web page content corresponding to a URL; and   at least one URL selection module in communication with the at least one URL data store and the at least one interaction data store;   wherein the at least one URL selection module selects at least one URL from the at least one URL data store that has interaction data in the at least one interaction data store; and   wherein the at least one URL selection module provides the at least one selected URL to the at least one download module.   
     
     
         14 . The web crawler according to  claim 13  further comprising:
 at least one content interest data store that stores an attention ranking of one or more content elements of a web page; and   at least one web page data store;   wherein the download module is configured to:
 utilize the at least one content interest data store to determine a content interest score of one or more content elements of a downloaded web page; and 
 store the one or more content elements of the downloaded web page in the at least one web page data store dependent on the respective content interest score. 
   
     
     
         15 . The web crawler according to  claim 14  wherein the at least one web page data store is configured to group content elements of a plurality of web pages according to their content interest score. 
     
     
         16 . The web crawler according to  claim 13  comprising an event server that:
 receives at least one event stream generated during an interaction with a web page on a client browser;   analyzes the event stream; and   stores the analyzed event stream in the interaction data store.   
     
     
         17 . The web crawler according to  claim 16  wherein the event server analyzes the at least one event stream to determine an event generator type of the event stream. 
     
     
         18 . The web crawler according to  claim 13  wherein the at least one interaction data store receives interaction data from an event server. 
     
     
         19 . A computer-readable medium comprising computer-executable instructions for execution by a processor, that, when executed, cause the processor to:
 select a Uniform Resource Locator (URL) from a URL data store;   look up the selected URL in an interaction data store to determine if interaction data exists for the selected URL in the interaction data store; and   if interaction data exists for the selected URL, provide the selected URL to a download module.   
     
     
         20 . The computer readable medium according to  claim 19  comprising computer-executable instructions for execution by the processor, that, when executed, cause the processor to:
 select a plurality of URLs from the URL data store that have corresponding interaction data in the interaction data store;   rank the plurality of URLs according to out-click event data of the interaction data;   provide at least one of the plurality of URLs to the download module;
 wherein the at least one URL provided to the download module is provided depending on the rank.

Join the waitlist — get patent alerts

Track US2010287152A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.